Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Building CLIA for Resource-Scarce African Languages

Domain:

natural language processing

Record type:

softwarepaper
Creator:
KulVas
Publisher:
IGI
Host:
Since most of the existing major search engines and commercial Information Retrieval (IR) systems are primarily designed for well-resourced European and Asian languages, they have paid little attention to the development of Cross-Language Information Access (CLIA) technologies for resource-scarce African languages. This paper presents the authors' experience in building CLIA for indigenous African languages, with a special focus on the development and evaluation of Oromo-English-CLIR. The authors have adopted a knowledge-based query translation approach to design and implement their initial Oromo-English CLIR (OMEN-CLIR). Apart from designing and building the first OMEN-CLIR from scratch, another major contribution of this study is assessing the performance of the proposed retrieval system at one of the well-recognized international Cross-Language Evaluation Forums like the CLEF campaign. The overall performance of OMEN-CLIR was found to be very promising and encouraging, given the limited amount of linguistic resources available for severely under-resourced African languages like Afaan Oromo.

Visit

doi.org

Tasks

information retrievalmachine translation

Languages

OromoOromo, Borana-Arsi-Guji

Similar

NLP Web Services for Resource-Scarce LanguagesDeveloping Core Technologies for Resource-Scarce Nguni LanguagesViability of Neural Networks for Core Technologies for Resource-Scarce LanguagesA word‐level language identification strategy for resource‐scarce languagesBuilding Corpora for Low-Resource Kenyan LanguagesA deep learning based multilingual hate speech detection for resource scarce languages

NLP Web Services for Resource-Scarce Languages

In this paper, we present a project where existing text-based core technologies were ported to Java-based web services from various architectures. These technologies were developed over a period of eight years through various government funded projects for 10 resou

Developing Core Technologies for Resource-Scarce Nguni Languages

The creation of linguistic resources is crucial to the continued growth of research and development

Viability of Neural Networks for Core Technologies for Resource-Scarce Languages

In this paper, the viability of neural network implementations of core technologies (the focus of th

A word‐level language identification strategy for resource‐scarce languages

ABSTRACT This study is based on the premise that it is possible to train compute

Building Corpora for Low-Resource Kenyan Languages

Natural Language Processing is a crucial frontier in artificial intelligence, with broad application

A deep learning based multilingual hate speech detection for resource scarce languages

Over the last decade, the increased use of social media has led to an increase in hateful activities