Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation

Domaine:

natural language processing

Type de record:

paper
In this paper, we introduce XGLUE, a new benchmark dataset that can be used to train large-scale cross-lingual pre-trained models using multilingual and bilingual corpora and evaluate their performance across a diverse set of cross-lingual tasks. Comparing to GLUE(Wang et al., 2019), which is labeled in English for natural language understanding tasks only, XGLUE has two main advantages: (1) it provides 11 diversified tasks that cover both natural language understanding and generation scenarios; (2) for each task, it provides labeled data in multiple languages. We extend a recent cross-lingual pre-trained model Unicoder(Huang et al., 2019) to cover both understanding and generation tasks, which is evaluated on XGLUE as a strong baseline. We also evaluate the base versions (12-layer) of Multilingual BERT, XLM and XLM-R for comparison

Visit

arxiv.orgaclanthology.org

Connected records

dataset

Tasks

topic classificationparsingquestion answeringsummarizationnamed entity recognition

Languages

Swahili

Licenses

Similaires

Pre-training Data Size Variation in mT5 and Cross-Lingual Transfer Performance on the XTREME-R BenchmarkUC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-trainingArtificial Code-Switching in Pre-Training for Cross-Lingual Retrieval RobustnessCross-Lingual Dialogue Dataset Creation via Outline-Based GenerationTARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language UnderstandingFew-Shot Cross-Lingual Stance Detection with Sentiment-Based Pre-Training

Pre-training Data Size Variation in mT5 and Cross-Lingual Transfer Performance on the XTREME-R Benchmark

Multi-lingual language models (LM), such as mBERT, XLM-R, mT5, mBART, have been remarkably successfu

UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training

Vision-and-language pre-training has achieved impressive success in learning multimodal representati

Artificial Code-Switching in Pre-Training for Cross-Lingual Retrieval Robustness

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Cross-Lingual Dialogue Dataset Creation via Outline-Based Generation

Multilingual task-oriented dialogue (ToD) facilitates access to services and information for many (c

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding

Few-Shot Cross-Lingual Stance Detection with Sentiment-Based Pre-Training

The goal of stance detection is to determine the viewpoint expressed in a piece of text towards a ta