Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
IsbAkhHajHus
Host:avatar
Being able to thoroughly assess massive multi-task language understanding (MMLU) capabilities is essential for advancing the applicability of multilingual language models. However, preparing such benchmarks in high quality native language is often costly and therefore limits the representativeness of evaluation datasets. While recent efforts focused on building more inclusive MMLU benchmarks, these are conventionally built using machine translation from high-resource languages, which may introduce errors and fail to account for the linguistic and cultural intricacies of the target languages. In this paper, we address the lack of native language MMLU benchmark especially in the under-represented Turkic language family with distinct morphosyntactic and cultural characteristics. We propose two benchmarks for Turkic language MMLU: TUMLU is a comprehensive, multilingual, and natively developed language understanding benchmark specifically designed for Turkic languages. It consists of middle- and high-school level questions spanning 11 academic subjects in Azerbaijani, Crimean Tatar, Karakalpak, Kazakh, Tatar, Turkish, Uyghur, and Uzbek. We also present TUMLU-mini, a more concise, balanced, and manually verified subset of the dataset. Using this dataset, we systematically evaluate a diverse range of open and proprietary multilingual large language models (LLMs), including Claude, Gemini, GPT, and LLaMA, offering an in-depth analysis of their performance across different languages, subjects, and alphabets. To promote further research and development in multilingual language understanding, we release TUMLU-mini and all corresponding evaluation scripts. Accepted to ACL 2025, Main Conference

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similar

Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language UnderstandingTARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language UnderstandingPolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and DialectsUnified Framework for Spoken Language Understanding and Summarization in Task-Based Human Dialog processingRevisiting non-English Text Simplification: A Unified Multilingual BenchmarkA Unified Orthography for Bantu Languages of Kenya

Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding

Spoken language understanding (SLU) is indispensable for half of all living languages that lack a fo

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evalua

Unified Framework for Spoken Language Understanding and Summarization in Task-Based Human Dialog processing

International audience Dialogue summarization aims to create a concise and coherent o

Revisiting non-English Text Simplification: A Unified Multilingual Benchmark

Recent advancements in high-quality, large-scale English resources have pushed the frontier of Engli

A Unified Orthography for Bantu Languages of Kenya