Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Linguini: A benchmark for language-agnostic linguistic reasoning

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
SánAlaRopSte
Hôte:avatar
We propose a new benchmark to measure a language model's linguistic reasoning skills without relying on pre-existing language-specific knowledge. The test covers 894 questions grouped in 160 problems across 75 (mostly) extremely low-resource languages, extracted from the International Linguistic Olympiad corpus. To attain high accuracy on this benchmark, models don't need previous knowledge of the tested language, as all the information needed to solve the linguistic puzzle is presented in the context. We find that, while all analyzed models rank below 25% accuracy, there is a significant gap between open and closed models, with the best-performing proprietary model at 24.05% and the best-performing open model at 8.84%.

Visit

arxiv.org

Tags

Computation and Language

Similaires

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning EvaluationModeLing: A Novel Dataset for Testing Linguistic Reasoning in Language ModelsLINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct LanguagesMultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial ReasoningUrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in UrduAraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practic

ModeLing: A Novel Dataset for Testing Linguistic Reasoning in Language Models

Large language models (LLMs) perform well on (at least) some evaluations of both few-shot multilingu

LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages

In this paper, we present the LingOly benchmark, a novel benchmark for advanced reasoning abilities

MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

We present AraLingBench: a fully human annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). The benchmark spans five core categories: grammar, morphology, spelling, reading comprehension, and syntax, through 150 ex