Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
DoğLiaBlaPra
Publisher:
arXiv
Host:avatar
LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English. There are now attempts to extend LLM-as-a-Judge to multilingual settings including low-resource languages. However, LLMs have limited proficiency in low-resource languages, and there is often no adequate human validation in these settings. To highlight the scope of the problem and current practices, we explore the use of LLM-as-a-Judge evaluators in ACL Anthology papers focusing on multilingual settings and low-resource languages across a diverse set of tasks. Out of 650 papers mentioning LLM-as-a-judge, only 33 of them focus on low-resource or multilingual settings. Our in-depth analysis of these papers indicates inconsistent evaluation outcomes, a tendency to overtrust LLM judgments in multilingual settings, and the widespread reliance on a single judge model per study. To help the NLP community further, we conclude with recommendations about how to use LLM-as-a-Judge in multilingual and low-resource settings. Under Review

Visit

doi.orgarxiv.org

Tags

Computation and Language (cs.CL)Artificial Intelligence (cs.AI)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource LanguagesToward Robust Multilingual Adaptation of LLMs for Low-Resource LanguagesMultilingual jailbreaking of LLMs using low-resource languagesBetter as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual ClassificationAmharic LLaMA and LLaVA: Multimodal LLMs for Low Resource LanguagesAdapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages

Multimodal LLMs are evolving from vision-language to tri-modality that see, hear, and read, yet pipe

Toward Robust Multilingual Adaptation of LLMs for Low-Resource Languages

Large language models (LLMs) continue to struggle with low-resource languages, primarily due to limi

Multilingual jailbreaking of LLMs using low-resource languages

Large Language Models (LLMs) remain vulnerable to jailbreak attempts that circumvent safety guardrai

Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification

Large Language Models (LLMs) have demonstrated remarkable multilingual capabilities, making them pro

Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages

Large Language Models (LLMs) like GPT-4 and LLaMA have shown incredible proficiency at natural langu

Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters

This paper explores the integration of graph knowledge from linguistic ontologies into multilingual