Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
ArčKleRobDob
Hôte:avatar
LLMs are routinely evaluated on language use, yet their explicit knowledge about linguistic structure remains poorly understood. Existing linguistic benchmarks focus on narrow phenomena, emphasize high-resource languages, and rarely test metalinguistic knowledge - explicit reasoning about language structure. We present a multilingual evaluation of metalinguistic knowledge in LLMs, based on the World Atlas of Language Structures (WALS), documenting 192 linguistic features across 2,660 languages. We convert WALS features into natural-language multiple-choice questions and evaluate models across documented languages. Using accuracy and macro F1, and comparing to chance and majority-class baselines, we assess performance and analyse variation across linguistic domains and language-related factors. Results show limited metalinguistic knowledge: GPT-4o performs best but achieves moderate accuracy (0.367), while open-source models lag. Although all models perform above chance, they fail to outperform the majority-class baseline, suggesting they capture broad cross-linguistic patterns but lack fine-grained distinctions. Performance varies by domain, partly reflecting differences in online visibility. At the language level, accuracy correlates with digital language status: languages with greater digital presence and resources are evaluated more accurately, while low-resource languages perform worse. Analysis of predictive factors confirms that resource-related indicators (Wikipedia size, corpus availability) are more informative than geographic, genealogical, or sociolinguistic factors. Overall, LLM metalinguistic knowledge appears fragmented and shaped mainly by data availability, rather than broadly generalizable grammatical competence. We release the benchmark as an open-source dataset to support evaluation across languages and encourage greater global linguistic diversity in future LLMs.

Visit

arxiv.org

Tags

Computation and Language

Similaires

Uncovering inequalities in new knowledge learning by large language models across different languagesMEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and TasksFrom Facts to Folklore: Evaluating Large Language Models on Bengali Cultural KnowledgeEvaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"Evaluating the Usage of African-American Vernacular English in Large Language ModelsSambaLingo: Teaching Large Language Models New Languages

Uncovering inequalities in new knowledge learning by large language models across different languages

As large language models (LLMs) gradually become integral tools for problem solving in daily life wo

MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks

There has been a surge in LLM evaluation research to understand LLM capabilities and limitations. Ho

From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge

Recent progress in NLP research has demonstrated remarkable capabilities of large language models (L

Evaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"

This paper evaluates the understanding and biases of large language models (LLMs) regarding

Evaluating the Usage of African-American Vernacular English in Large Language Models

In AI, most evaluations of natural language understanding tasks are conducted in standardized dialec

SambaLingo: Teaching Large Language Models New Languages

Despite the widespread availability of LLMs, there remains a substantial gap in their capabilities a