



A collection of encoder-based models trained on Southern African languages, utilizing language-specific subsets from a cleaned mC4 text corpus. The models are trained on the following languages and language combinations:
For each language or combination, the following models have been trained:
These models are used in Visser et al., "Insights Into Low-Resource Language Modelling: Improving Model Performances for South African Languages," Journal of Universal Computer Science, 2024, Insights into Low-Resource….