Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

BiniyamAjaw/amharic-corpus

Record type:

dataset
Creator:
Bin
Host:

Visit

huggingface.co

Languages

Amharic

Licenses

mit

Similar

BiniyamAjaw/amharic_tokenizerAmharic corpusContemporary Amharic Corpus: Automatically Morpho-Syntactically Tagged Amharic CorpusAmharic Pretraining CorpusNaolBM/amharic-corpusAmharic Speech Corpus

BiniyamAjaw/amharic_tokenizer

Amharic corpus

To build distributional semantic models, a large amount of text is required. These days, an enormous amount of texts are being generated continuously from different sources. As we want to build general-purpose semantic models, we collected datasets from different c

Contemporary Amharic Corpus: Automatically Morpho-Syntactically Tagged Amharic Corpus

We introduced the contemporary Amharic corpus, which is automatically tagged for morpho-syntactic information. Texts are collected from 25,199 documents from different domains and about 24 million orthographic words are tokenized. Since it is partly a web corpus, w

Amharic Pretraining Corpus

Amharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining

NaolBM/amharic-corpus

Amharic Speech Corpus

This is an Amharic speech corpus which is suitable for the development and evaluation of speech recognition and retrieval systems. The corpus contains 110 hours of speech data with syllable and grapheme-based transcriptions collected from public domain or resources