Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Linguistically enriched corpora for conjunctively written South African languages

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Puttkammer, MartinGaustad, Tanja
Éditeur:
Pienaar, WikusDu Toit, JacoGent, Sunny
Éditeur:
North-West University, Centre for Language Technology (CTexT)
Hôte:avatar
This resource contains linguistically annotated data for four official South African languages with a conjunctive orthography from the Nguni family (isiNdebele, isiXhosa, isiZulu and Siswati) as well as English. The data set is parallel for all five languages and the Nguni languages have been annotated for three different types of linguistic information: morphology, part-of-speech and lemmas. We have also included the protocols and tagsets used during annotation.

Visit

doi.org

Languages

NdebeleNdebeleNgwoSwatiXhosaZulu

Tags

Nguni languagesPOSMorphologyLemmaParallel data

Licenses

CC BY 4.0: https://creativecommons.org/licenses/by/4.0/