Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Construction of Amharic information retrieval resources and corpora

Domain:

natural language processing

Record type:

paperdataset
Creator:
YesMotAss
Editor:
AddSys
Publisher:
CCSDSpringer-Verlag
Host:avatar
International audience The development of information retrieval systems and natural language processing tools has been made possible for many natural languages because of the availability of natural language resources and corpora. Although Amharic is the working language of Ethiopia, it is still an under-resourced language. There are no adequate resources and corpora for Amharic ad-hoc retrieval evaluation to date. The existing ones are not publicly accessible and are not suitable for making scientific evaluation of information retrieval systems. To promote the development of Amharic adhoc retrieval, we build an ad-hoc retrieval test collection that consists of raw text, morphologically annotated stem-based and root-based corpora, a stopword list, stem-based and root-based lexicons, and WordNet-like resources. We also created word embeddings using the raw text and morphologically segmented forms of the corpora. When building these resources and corpora, we heavily consider the morphological characteristics of the language. The aim of this paper is to present these Amharic resources and corpora that we made available to the research community for information retrieval tasks. These resources and corpora are also evaluated experimentally and by linguists.

Visit

hal.science

Tasks

information retrieval

Languages

Amharic

Tags

EvaluationResourcesCorporaAmharic languageInformation retrieval[INFO.INFO-IR]Computer Science [cs]/Information Retrieval [cs.IR]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess