Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

rasyosef/amharic-sentences-corpus

Domain:

natural language processing

Record type:

dataset
Creator:
ras
Host:
This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining. GitHub uhh-lt/ethiopicmodels Dataset: Amharic corpus Paper: Introducing various Semanti… For citing this dataset, please use the following: @Article{fi13110275,

Visit

huggingface.co

Tasks

language modeling

Languages

Amharic

Similar

a3xrfgb/amharic-sentences-corpusmikeendale/gpt2-amharic-sentences-corpusrasyosef/amharic-sentimentrasyosef/bert-amharicrasyosef/gpt2-small-amharicrasyosef/amharic-llama-preferences

a3xrfgb/amharic-sentences-corpus

This 1.6 million Amharic sentences corpus reflects current Amharic usage as of December 20, 2025, an

mikeendale/gpt2-amharic-sentences-corpus

rasyosef/amharic-sentiment

This dataset contains 2781 cleaned Amharic Tweets, labeled as having either positive or negative sen

rasyosef/bert-amharic

BERT transformer models pretrained on amharic text # BERT Amharic This repo contains 4 BERT transf

rasyosef/gpt2-small-amharic

rasyosef/amharic-llama-preferences