AmharicDataRepo is a comprehensive repository containing datasets for pre-training and evaluating Natural Language Processing (NLP) models on Amharic text.
# AmharicDataRepo
AmharicDataRepo is a comprehensive repository containing datasets for pre-training and evaluating Natural Language Processing (NLP) models on Amharic text.
DESCRIPTION OF THE DATASET FOR PRE-TRAINING
The following table represents a collection of sources compiled to gather Amharic monolingual data, essential for pre-training language models. This compilation focuses on various domains, primarily news, with the inclusion of religious texts to provide a diverse linguistic dataset. Each source is identified by its type of content and the methods used for data extraction, such as sitemaps and the Wayback Machine, to ensure a comprehensive coverage of Amharic language content.
| Source | Domain | Number of Articles | Number of Sentences |
|-------------------------|------------|--------------------|------------------------------------|
| Zehabesha | News | 8,834 | Sitemap |
| DW | News | 68,160 | Sitemap & Wayback Machine |
| VOA | News | 178,961 | Sitemap & Wayback Machine |
| ESAT | News | 11,263 | Sitemap |
| Fana | News | 46,935 | Sitemap |
| Ethiopian Reporter | News | 23,051 | Sitemap |
| Amhara Media | News | 125 | Sitemap |
| Debteraw | News | 538 | Sitemap |
| Walta Information Center| News | 32,550 | Sitemap |
| EBC | News | 33,842 | Wayback Machine |
| BBC | News | 27,198 | Sitemap & Wayback Machine |
| ENA …