Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BiniyamAjaw/amharic-corpus

Type de record:

dataset
Créateur:
Bin
Hôte:

Visit

huggingface.co

Languages

Amharic

Licenses

mit

Similaires

BiniyamAjaw/amharic_tokenizerAmharic corpusContemporary Amharic Corpus: Automatically Morpho-Syntactically Tagged Amharic Corpusdagisky/Amharic-CorpusAmharic speech corpusNaolBM/amharic-corpus

BiniyamAjaw/amharic_tokenizer

Amharic corpus

To build distributional semantic models, a large amount of text is required. These days, an enormous amount of texts are being generated continuously from different sources. As we want to build general-purpose semantic models, we collected datasets from different c

Contemporary Amharic Corpus: Automatically Morpho-Syntactically Tagged Amharic Corpus

We introduced the contemporary Amharic corpus, which is automatically tagged for morpho-syntactic information. Texts are collected from 25,199 documents from different domains and about 24 million orthographic words are tokenized. Since it is partly a web corpus, w

dagisky/Amharic-Corpus

# Amharic Corpus Python based code to collect and preprocess Amharic Language Corpus ## Table

Amharic speech corpus

NaolBM/amharic-corpus