Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

The Development of Pre-processing Tools and Pre-trained Embedding Models for Amharic

Domain:

natural language processing

Record type:

datasetsoftwaremodel
Creator:
AssAyeBelYimam, Seid Muhie
Publisher:
Und
Host:avatar
Amharic is the second most spoken Semitic language after Arabic and serves as the official working language of Ethiopia. While Amharic NLP research is getting wider attentions recently, the main bottleneck is that the resources and related tools are not publicly released, which makes it still a low-resource language. Due to this reason, we observe that different researchers try to repeat the same NLP research again and again. In this work, we investigate the existing approach in Amharic NLP and take the first step to publicly release tools, datasets, and models to advance Amharic NLP research. We build Python-based preprocessing tools for Amharic (tokenizer, sentence segmenter, and text cleaner) that can easily be used and integrated for the development of NLP applications. Furthermore, we compiled the first moderately large-scale Amharic text corpus (6.8m sentences) along with the word2Vec, fastText, RoBERTa, and FLAIR embeddings models. Finally, we compile benchmark datasets and build classification models for the named entity recognition task.

Visit

doi.orgunderline.io

Tasks

embeddingsinformation extractionnamed entity recognitionsentence segmentation

Languages

Amharic

Tags

Natural Language ProcessingMachine LearningMachine Learning and Data MiningComputational LinguisticsLanguage ModelsNamed Entity Recognition