Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Annotation protocol for the Corpus of spoken isiXhosa

Domain:

natural language processing
Creator:
BloSlater, OnelisaZah
Publisher:
Zenodo
Host:avatar
This document lists the abbreviations used in the morphological annotation and part of speech tagging of the Corpus of spoken isiXhosa, and forms the result of years of work on this corpus. 

Visit

doi.orgzenodo.org

Tasks

part of speech tagging

Languages

Xhosa

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

xhosa Corpus of spoken isiXhosaSpoken Tunisian Arabic Corpus “STAC”: Transcription and AnnotationHuman sentiment annotation for isiXhosa and isiZulu.Morphologically annotated corpus for isiXhosaIsixhosa Ner CorpusMonolingual isiXhosa corpus

xhosa Corpus of spoken isiXhosa

The Corpus of Spoken isiXhosa The Corpus of Spoken isiXhosa consists of transcribed and annotated

Spoken Tunisian Arabic Corpus “STAC”: Transcription and Annotation

Human sentiment annotation for isiXhosa and isiZulu.

Human sentiment annotation for isiXhosa and isiZulu.

Morphologically annotated corpus for isiXhosa

NCHLT corpus of morphologically annotated tokens in isiXhosa converted to the tags used during phase

Isixhosa Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Monolingual isiXhosa corpus

Monolingual corpus for isiXhosa. The data is given as a single UTF-8 text file, with each segment on