Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

South African multilingual lexicons

Domaine:

natural language processing

Type de record:

dataset
Créateur:
ThaVukAbi
Hôte:avatar

This dataset contains a list of paired words for South Africa languages: Sepedi (nso), Sesotho(st), Tshivenda(ven), Xitsonga(tso), Setswana(tsn), IsiXhosa(xho), Isizulu(zul), Afrikaans(af), Isiswati(ssw), IsiNdebele(nr), and English(en). The paired words are stored in a json file with keys dict_keys(['en-af', 'en-zul', 'en-xho', 'en-ssw', 'en-nr', 'en-nso', 'en-tsn', 'en-st', 'en-ven', 'en-tso']) for ease of use and accessibility. For each key (E.g en-xho) retrieves a list of paired words between English (en) and IsiXhosa(xho).

Visit

figshare.com

Languages

AfrikaansNdebeleNdebeleSetswanaSotho, NorthernSotho, SouthernTsoTsongaXhosaZulu

Tags

Natural language processingBilingual lexiconMultilingual lexiconsSouth African languages parallel terminology listTerminology listLinguisticsMultilingualismSDG 4 Quality education

Licenses

CC BY-SA 4.0

Similaires

Towards machine-readable lexicons for South African Bantu languagesAn International Bibliography of African Lexicons.Biri lexiconsThe Vuk'uzenzele South African Multilingual CorpusThe Vuk'uzenzele South African Multilingual CorpusThe Gov South African Multilingual Corpus

Towards machine-readable lexicons for South African Bantu languages

Lexical information for South African Bantu languages is not readily available in the form of machine-readable lexicons. At present the availability of lexical information is restricted to a variety of paper dictionaries. These dictionaries display considerable div

An International Bibliography of African Lexicons.

Biri lexicons

The Vuk'uzenzele South African Multilingual Corpus

The dataset contains editions from the South African government magazine Vuk'uzenzele. Data was scraped from PDFs that have been placed in the data/raw folder. The PDFS were obtatined from the Vuk'uzenzele website (https://www.vukuzenzele.gov.za/). The datasets co

The Vuk'uzenzele South African Multilingual Corpus

Github: https://github.com/dsfsi/vukuzenzele-nlp/ Zenodo: Arxiv Preprint: Give Feedback 📑: DSFSI Res

The Gov South African Multilingual Corpus

The data set contains cabinet statements from the South African government, maintained by the Govern