Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Addressing Name Variation in Low-Resource Amharic Corpora: A Phonetic Normalization Strategy for Information Retrieval

Domaine:

natural language processing

Type de record:

model
Créateur:
NigHayAye
Éditeur:
Elsevier BV
Hôte:
For low-resource Amharic corpora, information retrieval faces significant challenges stemming from the language's orthographic redundancy, complex morphology, and inconsistent name representations. Variations in homophonic characters and complex affixation often lead to low recall in search systems. This research presents and evaluates the Robust Amharic Information Retrieval (RAIR) model, utilizing a symmetric normalization engine to standardize query and document formats through a three-tier method of orthographic collapsing, light stemming, and phonetic encoding. Experimental results derived from a large set of 144,201 news articles demonstrate an impressive 114% improvement in recall, rising from an original value of 0.43 to 0.92. The comparative analysis shows that the RAIR model outperforms conventional rule-based benchmarks and recent 2AIRTC results, establishing a new performance benchmark in low-resource Ethiopic information retrieval. The study highlights symmetric normalization as an essential requirement for contemporary Amharic search systems, effectively balancing the need for accurate results with the demand for comprehensive document retrieval.

Visit

doi.org

Tasks

information retrievaltext normalization

Languages

AmharicGeez

Similaires

Construction of Amharic information retrieval resources and corporaWord-based Probabilistic Phonetic Retrieval for Low-resource Spoken Term DetectionText Normalization for Low Resource LanguagesMorphology-Aware Retrieval for Low-Resource Environments: Advancing Information Retrieval for Shona LanguageInformation Retrieval Strategies for Accessing African Audio CorporaDense Retrieval for Low Resource Languages -- the Case of Amharic Language

Construction of Amharic information retrieval resources and corpora

International audience The development of information retrieval systems and natural l

Word-based Probabilistic Phonetic Retrieval for Low-resource Spoken Term Detection

Two problems make Spoken Term Detection (STD) particularly challenging under low-resource conditi

Text Normalization for Low Resource Languages

This repository contains code related to the Google open source internship project Text Normalization for Low Resource Languages.

Morphology-Aware Retrieval for Low-Resource Environments: Advancing Information Retrieval for Shona Language

Information Retrieval Strategies for Accessing African Audio Corpora

International audience In this paper we present a first approach to access African or

Dense Retrieval for Low Resource Languages -- the Case of Amharic Language

This paper reports some difficulties and some results when using dense retrievers on Amharic, one of