Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.
Monolingual corpus for SiSwati. The data is given as a single UTF-8 text file, with each segment on
This is the siSwati language part of the NCHLT Speech Corpus of the South African languages. Languag
Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems.
Orthographic and phonemically aligned transcriptions