Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama

Domain:

natural language processing

Record type:

paperdatasetproject
Creator:
EtoLu,KarKan
Host:avatar
As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are trained on multilingual data, but their reasoning abilities are mainly evaluated using English datasets. Hence, robust evaluation frameworks are needed using high-quality non-English datasets, especially low-resource languages (LRLs). This study evaluates eight state-of-the-art (SOTA) LLMs on Latvian and Giriama using a Massive Multitask Language Understanding (MMLU) subset curated with native speakers for linguistic and cultural relevance. Giriama is benchmarked for the first time. Our evaluation shows that OpenAI's o1 model outperforms others across all languages, scoring 92.8% in English, 88.8% in Latvian, and 70.8% in Giriama on 0-shot tasks. Mistral-large (35.6%) and Llama-70B IT (41%) have weak performance, on both Latvian and Giriama. Our results underscore the need for localized benchmarks and human evaluations in advancing cultural AI contextualization. Accepted at NoDaLiDa/Baltic-HLT 2025. hdl.handle.net

Visit

arxiv.org

Tasks

question answering

Languages

Kigiryama

Tags

Computation and Language

Similar

LLM Reporting of Official Statistics: Benchmarking the Accuracy and Reliability of Six Commercial LLMs from Frontier AI Labs Across 126 UNICEF-Published Child IndicatorsToxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource LanguagesBenchmarking Vision Language Models for Cultural UnderstandingVLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource LanguagesAfrican Languages - OpenAI MMLU2A2I-R/MMLU-Darija

LLM Reporting of Official Statistics: Benchmarking the Accuracy and Reliability of Six Commercial LLMs from Frontier AI Labs Across 126 UNICEF-Published Child Indicators

This study benchmarks two complementary dimensions of large-language-model (LLM) behaviour when aske

Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

The advancement of Large Language Models (LLMs) has transformed natural language processing; however

Benchmarking Vision Language Models for Cultural Understanding

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLM

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evalu

African Languages - OpenAI MMLU

This is a filtered subset of the openai/MMMLU dataset. This dataset only included mathematics (abstr

2A2I-R/MMLU-Darija