Language Adaptation Corpus for 17 African languages, English, French, and Arabic.
We used this corpus to train the following pre-trained language models:
If you use this corpus, please cite the MAFAND paper and mC4 paper.