Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Development of Pre-processing Tools and Pre-trained Embedding Models for Amharic

Domaine:

natural language processing

Type de record:

datasetsoftwaremodel
Créateur:
AssAyeBelYimam, Seid Muhie
Éditeur:
Und
Hôte:avatar
Amharic is the second most spoken Semitic language after Arabic and serves as the official working language of Ethiopia. While Amharic NLP research is getting wider attentions recently, the main bottleneck is that the resources and related tools are not publicly released, which makes it still a low-resource language. Due to this reason, we observe that different researchers try to repeat the same NLP research again and again. In this work, we investigate the existing approach in Amharic NLP and take the first step to publicly release tools, datasets, and models to advance Amharic NLP research. We build Python-based preprocessing tools for Amharic (tokenizer, sentence segmenter, and text cleaner) that can easily be used and integrated for the development of NLP applications. Furthermore, we compiled the first moderately large-scale Amharic text corpus (6.8m sentences) along with the word2Vec, fastText, RoBERTa, and FLAIR embeddings models. Finally, we compile benchmark datasets and build classification models for the named entity recognition task.

Visit

doi.orgunderline.io

Tasks

embeddingsinformation extractionnamed entity recognitionsentence segmentation

Languages

Amharic

Tags

Natural Language ProcessingMachine LearningMachine Learning and Data MiningComputational LinguisticsLanguage ModelsNamed Entity Recognition

Similaires

Dialect-Based Noisy Speech Dataset, Pre-Processing Tools, and Recognition Models for AmharicImproving Pre-trained Segmentation Models using Post-ProcessingImproving Pre-trained Adult Glioma Segmentation Models Using only Post-processing TechniquesPre-trained models: Galician, Iban, SetswanaArabic Extractive Summarization Using Pre-Trained ModelsMikyas1/amharic-etv-corpus-pre-processing

Dialect-Based Noisy Speech Dataset, Pre-Processing Tools, and Recognition Models for Amharic

Improving Pre-trained Segmentation Models using Post-Processing

Gliomas are the most common malignant brain tumors in adults and are among the most lethal. Despite

Improving Pre-trained Adult Glioma Segmentation Models Using only Post-processing Techniques

Gliomas are the most common malignant brain tumors in adults and are among the most lethal. Despite

Pre-trained models: Galician, Iban, Setswana

wav2vec 2.0 XLSR-128 models (with and without adaptation via continued pre-training)

Arabic Extractive Summarization Using Pre-Trained Models

Automatic Text Summarization (ATS) is a crucial area of study in Natural Language Processing (NLP) d

Mikyas1/amharic-etv-corpus-pre-processing

### TO GET TOTAL LINES OF A TEXT FILE - wc -l ### RESOURCES FOR CORPUS CLEANING AND PRE-PROCESSING