Logo Lanfrica

Document Classification for the Under-resourced Amharic Language

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AssWol
Éditeur:
Und
Hôte:avatar
NLP is severely hampered by a scarcity of digital resources. This is especially true for Amharic, a language with few resources but a rich morphology. In response, a total of 67,739 Amharic news documents from 8 different categories are gathered from web sources. A baseline document categorization experiment is carried out to validate the usability of the obtained corpora from various domains.