Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning

Domain:

natural language processing

Record type:

paperdataset
Creator:
PonGlaMajLiu
Host:avatar
In order to simulate human language capacity, natural language processing systems must be able to reason about the dynamics of everyday situations, including their possible causes and effects. Moreover, they should be able to generalise the acquired world knowledge to new languages, modulo cultural differences. Advances in machine reasoning and cross-lingual transfer depend on the availability of challenging evaluation benchmarks. Motivated by both demands, we introduce Cross-lingual Choice of Plausible Alternatives (XCOPA), a typologically diverse multilingual dataset for causal commonsense reasoning in 11 languages, which includes resource-poor languages like Eastern Apurímac Quechua and Haitian Creole. We evaluate a range of state-of-the-art models on this novel dataset, revealing that the performance of current methods based on multilingual pretraining and zero-shot fine-tuning falls short compared to translation-based transfer. Finally, we propose strategies to adapt multilingual models to out-of-sample resource-lean languages where only a small corpus or a bilingual dictionary is available, and report substantial improvements over the random baseline. The XCOPA dataset is freely available at github.com.

Visit

arxiv.org

Tasks

commonsense reasoningtransfer learning

Tags

Computation and Language

Similar

Fikira Dataset | A Multilingual Reasoning Dataset for African LanguagesCommon Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense ReasoningCommonsense Reasoning in Arab CultureCureMed-Bench: Multilingual Medical Reasoning DatasetA Deep Learning-Based Bengali Visual Commonsense Reasoning SystemBroaden the Vision: Geo-Diverse Visual Commonsense Reasoning

Fikira Dataset | A Multilingual Reasoning Dataset for African Languages

Fikira (Swahili for "thinking/reasoning") is a multilingual reasoning dataset for African languages, developed by Vambo AI. This dataset contains 50,000 reasoning examples across 10 African languages, synthetically generated as part of ongoing experiments at Vambo

Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense Reasoning

Commonsense reasoning research has so far been limited to English. We aim to evaluate and improve popular multilingual language models (ML-LMs) to help advance commonsense reasoning (CSR) beyond English. We collect the Mickey Corpus, consisting of 561k sentences in

Commonsense Reasoning in Arab Culture

Despite progress in Arabic large language models, such as Jais and AceGPT, their evaluation on commo

CureMed-Bench: Multilingual Medical Reasoning Dataset

CUREMED-BENCH is a multilingual medical reasoning benchmark dataset, designed for evaluating and fin

A Deep Learning-Based Bengali Visual Commonsense Reasoning System

Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning

Commonsense is defined as the knowledge on which everyone agrees. However, certain types of commonsense knowledge are correlated with culture and geographic locations and they are only shared locally. For example, the scenes of wedding ceremonies vary across region