Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TsedeniyaTemesgen/AmharicDataRepo

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Tse
Hôte:
AmharicDataRepo is a comprehensive repository containing datasets for pre-training and evaluating Natural Language Processing (NLP) models on Amharic text. # AmharicDataRepo AmharicDataRepo is a comprehensive repository containing datasets for pre-training and evaluating Natural Language Processing (NLP) models on Amharic text. DESCRIPTION OF THE DATASET FOR PRE-TRAINING The following table represents a collection of sources compiled to gather Amharic monolingual data, essential for pre-training language models. This compilation focuses on various domains, primarily news, with the inclusion of religious texts to provide a diverse linguistic dataset. Each source is identified by its type of content and the methods used for data extraction, such as sitemaps and the Wayback Machine, to ensure a comprehensive coverage of Amharic language content. | Source | Domain | Number of Articles | Number of Sentences | |-------------------------|------------|--------------------|------------------------------------| | Zehabesha | News | 8,834 | Sitemap | | DW | News | 68,160 | Sitemap & Wayback Machine | | VOA | News | 178,961 | Sitemap & Wayback Machine | | ESAT | News | 11,263 | Sitemap | | Fana | News | 46,935 | Sitemap | | Ethiopian Reporter | News | 23,051 | Sitemap | | Amhara Media | News | 125 | Sitemap | | Debteraw | News | 538 | Sitemap | | Walta Information Center| News | 32,550 | Sitemap | | EBC | News | 33,842 | Wayback Machine | | BBC | News | 27,198 | Sitemap & Wayback Machine | | ENA …

Visit

github.com

Tasks

language modeling

Languages

Amharic

Licenses

GPL-3.0