Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Amharic Pretraining Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
yor
Host:
Amharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic. You can load the dataset as follows from datasets import load_dataset ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus")

Visit

huggingface.co

Tasks

language modeling

Languages

Amharic

Licenses

apache-2.0

Similar

Lyte/darija-pretraining-corpusMoroccan Darija Pretraining Corpus for NanochatAmharic corpusManeno Yetu: Dynamic Corpus Construction and Pretraining for Swahili NLPContemporary Amharic Corpus: Automatically Morpho-Syntactically Tagged Amharic CorpusAdapting BERT and AgriBERT for agroecology: A small-corpus pretraining approach

Lyte/darija-pretraining-corpus

Moroccan Darija Pretraining Corpus for Nanochat

Parquet shards prepared for nanochat pretraining. 83 train shards 1 validation shard 82,738,910 tra

Amharic corpus

To build distributional semantic models, a large amount of text is required. These days, an enormous amount of texts are being generated continuously from different sources. As we want to build general-purpose semantic models, we collected datasets from different c

Maneno Yetu: Dynamic Corpus Construction and Pretraining for Swahili NLP

Contemporary Amharic Corpus: Automatically Morpho-Syntactically Tagged Amharic Corpus

We introduced the contemporary Amharic corpus, which is automatically tagged for morpho-syntactic information. Texts are collected from 25,199 documents from different domains and about 24 million orthographic words are tokenized. Since it is partly a web corpus, w

Adapting BERT and AgriBERT for agroecology: A small-corpus pretraining approach

Source Agritrop Cirad (https://agritrop.cirad.fr/614086/) * Autres projet