Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Multilingual Curse at the Retrieval Layer: Evidence from Amharic

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
AleMekde
Hôte:avatar
Multilingual retrieval increasingly underpins cross-lingual question answering and retrieval-augmented generation. Strong zero-shot scores on multilingual benchmarks are often taken as evidence that current encoders transfer reliably across many languages. We argue that this assumption breaks down for underrepresented, morphologically rich languages, and use Amharic as a diagnostic case. Under a shared passage retrieval protocol covering dense, late-interaction, learned sparse, and cross-encoder paradigms, we compare zero-shot multilingual retrievers, Amharic-fine-tuned multilingual retrievers, and monolingual Amharic retrievers. The strongest zero-shot multilingual retriever underperforms the strongest monolingual Amharic first-stage retriever by 23% relative MRR@10. Fine-tuning two recent multilingual embedding models on the same Amharic supervision yields 32-60% relative MRR@10 gains over zero-shot, but the best Amharic-fine-tuned multilingual model remains below the strongest monolingual Amharic retriever. These findings indicate that zero-shot multilingual retrieval is not a sufficient proxy for equitable information access in the LLM era: for underrepresented languages, retrieval must be evaluated and adapted in-language rather than inferred from aggregate multilingual benchmarks. To foster future research, we publicly release the dataset, codebase, and trained models at rasyosef/amharic-neural-ir. 10 pages, 4 tables. Accepted to the 1st Workshop on Multilinguality in the Era of Large Language Models (MeLLM) at ACL 2026

Visit

arxiv.org

Tasks

information retrieval

Languages

Amharic

Tags

Information RetrievalComputation and LanguageMachine LearningH.3.3; I.2.7

Similaires

Rethinking the Multilingual Reasoning Gap with Layer Swap2AIRTC: The Amharic Adhoc Information Retrieval Test CollectionThe effectiveness of stemming for information retrieval in AmharicAmharic-English Information RetrievalCultural Variations in the Curse of Knowledge: the Curse of Knowledge Bias in Children from a Nomadic Pastoralist Culture in KenyaDoes Human Capital Mitigate Resource Curse? Evidence in the Short- and Long-Run

Rethinking the Multilingual Reasoning Gap with Layer Swap

Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, ev

2AIRTC: The Amharic Adhoc Information Retrieval Test Collection

International audience Evaluation is highly important for designing, developing, and

The effectiveness of stemming for information retrieval in Amharic

Amharic is an example of a language with a very rich morphology, which means that systems for search

Amharic-English Information Retrieval

Cultural Variations in the Curse of Knowledge: the Curse of Knowledge Bias in Children from a Nomadic Pastoralist Culture in Kenya

Abstract We examined the universality of the curse of knowledge (i.e., the tendency to be biased by

Does Human Capital Mitigate Resource Curse? Evidence in the Short- and Long-Run

This paper examines the short-term and long-term relationships among natural resources, human capita