Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Core technologies for conjunctively written South African languages

Domain:

natural language processing

Record type:

datasetsoftware
Creator:
Du Toit, JacoPuttkammer, Martin
Editor:
Gent, SunnyGaustad, Tanja
Publisher:
North-West University, Centre for Language Technology (CTexT)
Host:avatar
During this SADiLaR funded project, enriched corpora for the four official South African languages with a conjunctive orthography, i.e. isiNdebele (NR), isiXhosa (XH), isiZulu (ZU), and Siswati (SS) was developed. The corpora consist of approximately 50,000 tokens, parallel on sentence level, with English as source language, for each language. Each language’s corpus was annotated on three levels, namely morphological analysis, part of speech and lemmatisation (see: repo.sadilar.org). Using the annotated data, 12 core technologies, i.e. morphological analysers, POS taggers and lemmatisers for each of the four languages were developed and packaged in a single graphical user interface (UI).

Visit

hdl.handle.net

Tasks

part of speech tagging

Languages

NdebeleNdebeleSwatiXhosaZulu

Tags

part of speechpart of speech taggingpart-of-speech taggingpart-of-speechlemmalemmatisationlemmatizationmorphologymorphological analysisconjunctive languages+1