Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

xhosa Corpus of spoken isiXhosa

Domain:

natural language processing

Record type:

dataset
Creator:
Spr
Publisher:
Spr
Host:avatar
The Corpus of Spoken isiXhosa The Corpus of Spoken isiXhosa consists of transcribed and annotated recordings of spoken Xhosa [xho]. The recordings have been made in the Eastern Cape in South Africa from 2015 onwards. The transcribed texts are annotated with morpheme-by-morpheme glosses, part-of-speech tags, and free English translations. The recordings and the annotations of Xhosa data have been made as part of three different research projects led by senior lecturer Eva-Marie Bloom Ström at the University of Gothenburg. All projects, including the ongoing ‘How do words get in order? The role of speaker-hearer interaction in languages of southern Africa’, were founded by the Swedish Research Council. The Corpus has been developed in collaboration with Språkbanken Text. A user guide and more extensive information about the corpus data can be found in the Corpus of Spoken isiXhosa Manual [PDF]. For more on annotation, preparation of data, and acknowledgements see: Bloom Ström, E.-M., Slater, O., Zahran, A., Berdicevskis, A., & Schumacher, A. (2023). Preparing a corpus of spoken Xhosa. Proceedings of the 2023 CLASP Conference on Learning with Small Data (LSD), 62–67. aclanthology.org For questions about the corpus: Eva-Marie Bloom Ström eva-marie.strom@gu.se If you notice any errors or inconsistencies in annotations, please report them to this email address. Main contributors: Eva-Marie Bloom Ström Senior Lecturer, University of Gothenburg Onelisa Slater MA, Rhodes University Aron Zahran PhD, Inalco/Llacan (CNRS) & Ghent University The Corpus of Spoken isiXhosa The Corpus of Spoken isiXhosa consists of transcribed and annotated recordings of spoken Xhosa [xho]. The recordings have been made in the Eastern Cape in South Africa from 2015 onwards. The transcribed texts are annotated with morpheme-by-morpheme glosses, part-of-speech tags, and free English translations. The recordings and the annotations of Xhosa data have been made as part of three different research projects led by senior lecturer Eva-Marie Bloom Ström at the University of Gothenburg. All projects, including the ongoing ‘How do words get in order? The role of speaker-hearer interaction in languages of southern Africa’, were founded by the Swedish Research Council. The Corpus has been developed in collaboration with Språkbanken Text. A user guide and more extensive information about the corpus data can be found in the Corpus of Spoken isiXhosa Manual [PDF]. For more on annotation, preparation of data, and acknowledgements see: Bloom Ström, E.-M., Slater, O., Zahran, A., Berdicevskis, A., & Schumacher, A. (2023). Preparing a corpus of spoken Xhosa. Proceedings of the 2023 CLASP Conference on Learning with Small Data (LSD), 62–67. aclanthology.org For questions about the corpus: Eva-Marie Bloom Ström eva-marie.strom@gu.se If you notice any errors or inconsistencies in annotations, please report them to this email address. Main contributors: Eva-Marie Bloom Ström Senior Lecturer, University of Gothenburg Onelisa Slater MA, Rhodes University Aron Zahran PhD, Inalco/Llacan (CNRS) & Ghent University

Visit

doi.orgsprakbanken.se

Languages

Xhosa

Tags

Language Technology (Computational Linguistics)

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Annotation protocol for the Corpus of spoken isiXhosaStarting with Xhosa English towards a spoken corpusEXPLORING THE USE OF A SPOKEN XHOSA CORPUS FOR DEVELOPING XHOSA ADDITIONAL LANGUAGE TEACHING MATERIALSThe use of actually in spoken Xhosa English: a corpus studyXhosa/isiXhosa speech developmentIsixhosa Ner Corpus

Annotation protocol for the Corpus of spoken isiXhosa

This document lists the abbreviations used in the morphological annotation and part of speech taggin

Starting with Xhosa English towards a spoken corpus

This paper describes the underlying motivation for the proposed structure and design of a corpus of

EXPLORING THE USE OF A SPOKEN XHOSA CORPUS FOR DEVELOPING XHOSA ADDITIONAL LANGUAGE TEACHING MATERIALS

The use of actually in spoken Xhosa English: a corpus study

Xhosa/isiXhosa speech development

Abstract isiXhosa is spoken in South Africa, and there are several varieties inc

Isixhosa Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.