Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

Domain:

natural language processing

Record type:

datasetpaper
The Cross-lingual Natural Language Inference (XNLI) corpus is a crowd-sourced collection of 5,000 test and 2,500 dev pairs for the MultiNLI corpus. The pairs are annotated with textual entailment and translated into 14 languages: French, Spanish, German, Greek, Bulgarian, Russian, Turkish, Arabic, Vietnamese, Thai, Chinese, Hindi, Swahili and Urdu. This results in 112.5k annotated pairs. Each premise can be associated with the corresponding hypothesis in the 15 languages, summing up to more than 1.5M combinations. The corpus is made to evaluate how to perform inference in any language (including low-resources ones like Swahili or Urdu) when only English NLI data is available at training time. One solution is cross-lingual sentence encoding, for which XNLI is an evaluation benchmark. The Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark is a benchmark for the evaluation of the cross-lingual generalization ability of pre-trained multilingual models. It covers 40 typologically diverse languages (spanning 12 language families) and includes nine tasks that collectively require reasoning about different levels of syntax and semantics. The languages in XTREME are selected to maximize language diversity, coverage in existing tasks, and availability of training data. Among these are many under-studied languages, such as the Dravidian languages Tamil (spoken in southern India, Sri Lanka, and Singapore), Telugu and Malayalam (spoken mainly in southern India), and the Niger-Congo languages Swahili and Yoruba, spoken in Africa.

Visit

github.com

Tasks

part of speech taggingnatural language inferencequestion answeringnamed entity recognitiontransfer learning

Languages

AfrikaansSwahiliYoruba

Licenses

Similar

XLNet's Zero-Shot Cross-Lingual Generalization with Intermediate-Task Training in XTREME-RMulti-source Intermediate-task Training for Low-resource XTREME Language GeneralizationScaling Multilingual Language Models for Low-Resource Cross-Lingual NER on the XTREME BenchmarkMulti-Task Intermediate Training for Zero-Shot Cross-Lingual Transfer in XTREME-RMultilingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer Performance on XTREME-RMVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching

XLNet's Zero-Shot Cross-Lingual Generalization with Intermediate-Task Training in XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Multi-source Intermediate-task Training for Low-resource XTREME Language Generalization

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Scaling Multilingual Language Models for Low-Resource Cross-Lingual NER on the XTREME Benchmark

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Multi-Task Intermediate Training for Zero-Shot Cross-Lingual Transfer in XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Multilingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer Performance on XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching

Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages. Conse