Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense Reasoning

Domain:

natural language processing

Record type:

paper
Commonsense reasoning research has so far been limited to English. We aim to evaluate and improve popular multilingual language models (ML-LMs) to help advance commonsense reasoning (CSR) beyond English. We collect the Mickey Corpus, consisting of 561k sentences in 11 different languages, which can be used for analyzing and improving ML-LMs. We propose Mickey Probe, a language-agnostic probing task for fairly evaluating the common sense of popular ML-LMs across different languages. In addition, we also create two new datasets, X-CSQA and X-CODAH, by translating their English versions to 15 other languages, so that we can evaluate popular ML-LMs for cross-lingual commonsense reasoning. To improve the performance beyond English, we propose a simple yet effective method -- multilingual contrastive pre-training (MCP). It significantly enhances sentence representations, yielding a large performance gain on both benchmarks.

Visit

arxiv.orginklab.usc.edu

Connected records

dataset

Tasks

question answeringcommonsense reasoning

Languages

Swahili

Licenses

Similar

Evaluating Multilingual Long-Context Models for Retrieval and ReasoningXCOPA: A Multilingual Dataset for Causal Commonsense ReasoningGlobal PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and CulturesImproving Multilingual Math Reasoning for African LanguagesGeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsGlobal PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures v0.1

Evaluating Multilingual Long-Context Models for Retrieval and Reasoning

Recent large language models (LLMs) demonstrate impressive capabilities in handling long contexts, s

XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning

In order to simulate human language capacity, natural language processing systems must be able to re

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (

Improving Multilingual Math Reasoning for African Languages

Researchers working on low-resource languages face persistent challenges due to limited data availab

GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models

Recent work has shown that Pre-trained Language Models (PLMs) store the relational knowledge learned

Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures v0.1

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning benchmark for over 100 lan