This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining.
GitHub uhh-lt/ethiopicmodels
Dataset: Amharic corpus
Paper: Introducing various Semanti…
For citing this dataset, please use the following:
@Article{fi13110275,