Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

anrilombard/mzansi-text

Domain:

natural language processing

Record type:

dataset
Creator:
anr
Host:
MzansiText is a curated multilingual pretraining corpus for all eleven official South African languages. Languages: af, en, nso, sot, ssw, tsn, tso, ven, xho, zul, nbl Schema: { "text": "string", "lang": "string" } This repository contains the raw train, validation, and test text splits used for the MzansiLM pretraining release. The token distribution table below matches the paper-reported corpus statistics.

Visit

huggingface.co

Tasks

language modeling

Languages

AfrikaansNdebeleSetswanaSotho, NorthernSotho, SouthernSwatiTsongaVendaXhosaZulu

Tags

pretrainingsouth-african-languagesmultilingualmzansitextlarge datasets from Lanfrica Insights

Licenses

apache-2.0

Similar

anrilombard/mzansi-text-tokenized

anrilombard/mzansi-text-tokenized

Ready-to-train tokenized version of MzansiText, chunked to a context length of 2048 tokens. Tokeniz