Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
DemBabHämAli
Éditeur:
arXiv
Hôte:avatar
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.

Visit

doi.org

Tags

Computation and Language (cs.CL)Artificial Intelligence (cs.AI)FOS: Computer and information sciences

Licenses

Creative Commons Attribution Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-sa/4.0/legalcode

Similaires

MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning EvaluationEnhancing Low-Resource Language Reasoning via High-Resource Language Feature TransferReasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced ContextsQuantized LLMs for Zero-Shot Cross-Lingual Reasoning in Low-Resource LanguagesAn Analysis of Prospective Teachers’ Mathematical Reasoning on Number ConceptsCombining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource Languages

MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation

Large language models have made substantial progress in mathematical reasoning. However, benchmark d

Enhancing Low-Resource Language Reasoning via High-Resource Language Feature Transfer

Large language models exhibit substantial performance variation across languages, even when solving

Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts

Sentiment analysis in low-resource, culturally nuanced contexts challenges conventional NLP approach

Quantized LLMs for Zero-Shot Cross-Lingual Reasoning in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

An Analysis of Prospective Teachers’ Mathematical Reasoning on Number Concepts

This paper presents and discusses the results of a case study that was carried out to understand the

Combining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource Languages

The contrast between the need for large amounts of data for current Natural Language Processing (NLP) techniques, and the lack thereof, is accentuated in the case of African languages, most of which are considered low-resource. To help circumvent this issue, we exp