Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Core technologies for conjunctively written South African languages

Domaine:

natural language processing

Type de record:

datasetsoftware
Créateur:
Du Toit, JacoPuttkammer, Martin
Éditeur:
Gent, SunnyGaustad, Tanja
Éditeur:
North-West University, Centre for Language Technology (CTexT)
Hôte:avatar
During this SADiLaR funded project, enriched corpora for the four official South African languages with a conjunctive orthography, i.e. isiNdebele (NR), isiXhosa (XH), isiZulu (ZU), and Siswati (SS) was developed. The corpora consist of approximately 50,000 tokens, parallel on sentence level, with English as source language, for each language. Each language’s corpus was annotated on three levels, namely morphological analysis, part of speech and lemmatisation (see: repo.sadilar.org). Using the annotated data, 12 core technologies, i.e. morphological analysers, POS taggers and lemmatisers for each of the four languages were developed and packaged in a single graphical user interface (UI).

Visit

hdl.handle.net

Tasks

part of speech tagging

Languages

NdebeleNdebeleSwatiXhosaZulu

Tags

part of speechpart of speech taggingpart-of-speech taggingpart-of-speechlemmalemmatisationlemmatizationmorphologymorphological analysisconjunctive languages+1