Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Encoder Language Models for Southern African Languages

Domain:

natural language processing

Record type:

model
Creator:
Visser, Ruan
Editor:
Grobler, TrienkoDunaiski, Marcel
Publisher:
Zenodo
Host:avatar

A collection of encoder-based models trained on Southern African languages, utilizing language-specific subsets from a cleaned mC4 text corpus. The models are trained on the following languages and language combinations:

  • Zulu
  • Xhosa
  • Swahili
  • Zulu + Xhosa (ZX)
  • Zulu + Xhosa + Swahili (ZXS)
  • Zulu + Xhosa + Swahili + Shona + Nyanja (All)

For each language or combination, the following models have been trained:

  • BERT-base
  • BERT-small
  • ConvBERT-base
  • ConvBERT-small
  • ELECTRA-base
  • ELECTRA-small

These models are used in Visser et al., "Insights Into Low-Resource Language Modelling: Improving Model Performances for South African Languages," Journal of Universal Computer Science, 2024, Insights into Low-Resource….

Visit

doi.org

Tasks

language modeling

Languages

ChichewaShonaSwahiliXhosaZulu

Licenses

info:eu-repo/semantics/restrictedAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode