Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TsedeniyaTemesgen/AmharicDataRepo

Domain:

natural language processing

Record type:

dataset
Creator:
Tse
Host:
AmharicDataRepo is a comprehensive repository containing datasets for pre-training and evaluating Natural Language Processing (NLP) models on Amharic text. # AmharicDataRepo AmharicDataRepo is a comprehensive repository containing datasets for pre-training and evaluating Natural Language Processing (NLP) models on Amharic text. DESCRIPTION OF THE DATASET FOR PRE-TRAINING The following table represents a collection of sources compiled to gather Amharic monolingual data, essential for pre-training language models. This compilation focuses on various domains, primarily news, with the inclusion of religious texts to provide a diverse linguistic dataset. Each source is identified by its type of content and the methods used for data extraction, such as sitemaps and the Wayback Machine, to ensure a comprehensive coverage of Amharic language content. | Source | Domain | Number of Articles | Number of Sentences | |-------------------------|------------|--------------------|------------------------------------| | Zehabesha | News | 8,834 | Sitemap | | DW | News | 68,160 | Sitemap & Wayback Machine | | VOA | News | 178,961 | Sitemap & Wayback Machine | | ESAT | News | 11,263 | Sitemap | | Fana | News | 46,935 | Sitemap | | Ethiopian Reporter | News | 23,051 | Sitemap | | Amhara Media | News | 125 | Sitemap | | Debteraw | News | 538 | Sitemap | | Walta Information Center| News | 32,550 | Sitemap | | EBC | News | 33,842 | Wayback Machine | | BBC | News | 27,198 | Sitemap & Wayback Machine | | ENA …

Visit

github.com

Tasks

language modeling

Languages

Amharic

Licenses

GPL-3.0