Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
BeaHelMayMag
Host:avatar
In this paper, we present the LingOly benchmark, a novel benchmark for advanced reasoning abilities in large language models. Using challenging Linguistic Olympiad puzzles, we evaluate (i) capabilities for in-context identification and generalisation of linguistic patterns in very low-resource or extinct languages, and (ii) abilities to follow complex task instructions. The LingOly benchmark covers more than 90 mostly low-resource languages, minimising issues of data contamination, and contains 1,133 problems across 6 formats and 5 levels of human difficulty. We assess performance with both direct accuracy and comparison to a no-context baseline to penalise memorisation. Scores from 11 state-of-the-art LLMs demonstrate the benchmark to be challenging, and models perform poorly on the higher difficulty problems. On harder problems, even the top model only achieved 38.7% accuracy, a 24.7% improvement over the no-context baseline. Large closed models typically outperform open models, and in general, the higher resource the language, the better the scores. These results indicate, in absence of memorisation, true multi-step out-of-domain reasoning remains a challenge for current language models. Oral presentation at NeurIPS 2024 Datasets and Benchmarks Track. 10 pages, 5 figures, 22 pages supplemental materials

Visit

arxiv.org

Tags

Computation and Language

Similar

LSR: Linguistic Safety Robustness Benchmark for Low-Resource West African LanguagesShadman19/llm-benchmark-low-resource-languagesLinguini: A benchmark for language-agnostic linguistic reasoningUNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?Linguistic Diversity in Intermediate Tasks and Zero-Shot F1 Scores on Low-Resource Languages within the XTREME BenchmarkLow Resource Neural Machine Translation: A Benchmark for Five African Languages

LSR: Linguistic Safety Robustness Benchmark for Low-Resource West African Languages

Safety alignment in large language models relies predominantly on English-language training data. Wh

Shadman19/llm-benchmark-low-resource-languages

Evaluating open-source LLMs on Bengali, Swahili and Tamil vs English baseline # 🌍 LLM Benchmark for

Linguini: A benchmark for language-agnostic linguistic reasoning

We propose a new benchmark to measure a language model's linguistic reasoning skills without relying

UNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?

Large language models (LLMs) have demonstrated potential in reasoning tasks, but their performance o

Linguistic Diversity in Intermediate Tasks and Zero-Shot F1 Scores on Low-Resource Languages within the XTREME Benchmark

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Low Resource Neural Machine Translation: A Benchmark for Five African Languages

Recent advents in Neural Machine Translation (NMT) have shown improvements in low-resource language (LRL) translation tasks. In this work, we benchmark NMT between English and five African LRL pairs (Swahili, Amharic, Tigrigna, Oromo, Somali [SATOS]). We collected