Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

Domaine:

natural language processing

Type de record:

paper
Créateur:
AdjEiselen, RoaldMit
Hôte:avatar
We investigate the translation quality of current large language models (LLMs) for English-to-Hausa and English-to-Fongbe - two typologically distinct West African languages from the Afroasiatic and Niger-Congo families respectively - and evaluate whether standard automatic metrics reliably reflect human judgment for these low-resource languages. We evaluate four models (GPT-4o Mini, Claude Sonnet 4, Gemini 2.5 Flash, and Qwen2.5-7B) at progressive scales (500 to 10,000 sentences) using automatic metrics (BLEU, chrF++, TER, COMET, BERTScore) validated against native-speaker judgment. Our results reveal three key findings. First, translation quality varies substantially by language: Hausa achieves acceptable quality (human scores 4.0-4.5/5) while Fongbe achieves poor quality (1.0-2.2/5), with a consistent 3x BLEU gap across all systems. Second, model rankings differ by language - Gemini leads for Fongbe while GPT-4o leads for Hausa by human evaluation - indicating that performance on one low-resource African language does not predict performance on another. Third, metric-human correlation varies dramatically: perfect rank correlation for Fongbe (rho=1.0) but weak correlation for Hausa (rho=0.5), where human evaluators preferred GPT-4o despite all automatic metrics ranking Claude first. We further show that neural metrics like BERTScore exhibit embedding collapse (within-language similarity >0.99) for both languages, limiting their ability to differentiate translation quality. Based on these findings, we recommend multi-metric evaluation for low-resource African languages, with particular caution when interpreting neural metrics. We establish that minimum sample sizes of n=2,500 sentences are required for stable system rankings, as smaller samples produced artifact findings that reversed at scale. 19 pages, 10 tables

Visit

arxiv.org

Tasks

machine translation

Languages

FonHausa

Tags

Computation and LanguageArtificial IntelligenceMachine LearningI.2.7

Similaires

Evaluating Quantized Large Language Models for Code Generation on Low-Resource Language BenchmarksMining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and FongbeEvaluating Large Language Models for Low-Resource Multilingual Machine Translation in the Medical DomainFrom Translation to Retrieval: Evaluating LLM-Based Information Retrieval for Hausa and FongbeAdaptive Machine Translation with Large Language ModelsInteractive Machine Translation with Large Language Models for Low-resource Languages

Evaluating Quantized Large Language Models for Code Generation on Low-Resource Language Benchmarks

Democratization of AI is an important topic within the broader topic of the digital divide. This iss

Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe

Large language models (LLMs) are trained on data contributed by low-resource language communities, y

Evaluating Large Language Models for Low-Resource Multilingual Machine Translation in the Medical Domain

This dissertation explores neural machine translation (NMT) in multilingual medical domain, with

From Translation to Retrieval: Evaluating LLM-Based Information Retrieval for Hausa and Fongbe

Adaptive Machine Translation with Large Language Models

Consistency is a key requirement of high-quality translation. It is especially important to adhere t

Interactive Machine Translation with Large Language Models for Low-resource Languages

Large language models (LLM) have been applied to machine translation with notable success. However,