Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama

Domaine:

natural language processing

Type de record:

paperdatasetproject
Créateur:
EtoLu,KarKan
Hôte:avatar
As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are trained on multilingual data, but their reasoning abilities are mainly evaluated using English datasets. Hence, robust evaluation frameworks are needed using high-quality non-English datasets, especially low-resource languages (LRLs). This study evaluates eight state-of-the-art (SOTA) LLMs on Latvian and Giriama using a Massive Multitask Language Understanding (MMLU) subset curated with native speakers for linguistic and cultural relevance. Giriama is benchmarked for the first time. Our evaluation shows that OpenAI's o1 model outperforms others across all languages, scoring 92.8% in English, 88.8% in Latvian, and 70.8% in Giriama on 0-shot tasks. Mistral-large (35.6%) and Llama-70B IT (41%) have weak performance, on both Latvian and Giriama. Our results underscore the need for localized benchmarks and human evaluations in advancing cultural AI contextualization. Accepted at NoDaLiDa/Baltic-HLT 2025. hdl.handle.net

Visit

arxiv.org

Tasks

question answering

Languages

Kigiryama

Tags

Computation and Language

Similaires

LLM Reporting of Official Statistics: Benchmarking the Accuracy and Reliability of Six Commercial LLMs from Frontier AI Labs Across 126 UNICEF-Published Child IndicatorsToxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource LanguagesBenchmarking Vision Language Models for Cultural UnderstandingVLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource LanguagesAfrican Languages - OpenAI MMLU2A2I-R/MMLU-Darija

LLM Reporting of Official Statistics: Benchmarking the Accuracy and Reliability of Six Commercial LLMs from Frontier AI Labs Across 126 UNICEF-Published Child Indicators

This study benchmarks two complementary dimensions of large-language-model (LLM) behaviour when aske

Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

The advancement of Large Language Models (LLMs) has transformed natural language processing; however

Benchmarking Vision Language Models for Cultural Understanding

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLM

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evalu

African Languages - OpenAI MMLU

This is a filtered subset of the openai/MMMLU dataset. This dataset only included mathematics (abstr

2A2I-R/MMLU-Darija