Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Addressing Name Variation in Low-Resource Amharic Corpora: A Phonetic Normalization Strategy for Information Retrieval

Domain:

natural language processing

Record type:

model
Creator:
NigHayAye
Publisher:
Elsevier BV
Host:
For low-resource Amharic corpora, information retrieval faces significant challenges stemming from the language's orthographic redundancy, complex morphology, and inconsistent name representations. Variations in homophonic characters and complex affixation often lead to low recall in search systems. This research presents and evaluates the Robust Amharic Information Retrieval (RAIR) model, utilizing a symmetric normalization engine to standardize query and document formats through a three-tier method of orthographic collapsing, light stemming, and phonetic encoding. Experimental results derived from a large set of 144,201 news articles demonstrate an impressive 114% improvement in recall, rising from an original value of 0.43 to 0.92. The comparative analysis shows that the RAIR model outperforms conventional rule-based benchmarks and recent 2AIRTC results, establishing a new performance benchmark in low-resource Ethiopic information retrieval. The study highlights symmetric normalization as an essential requirement for contemporary Amharic search systems, effectively balancing the need for accurate results with the demand for comprehensive document retrieval.

Visit

doi.org

Tasks

information retrievaltext normalization

Languages

AmharicGeez