Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.
Monolingual corpus for SiSwati. The data is given as a single UTF-8 text file, with each segment on
This is the siSwati language part of the NCHLT Speech Corpus of the South African languages. Languag
Orthographic and phonemically aligned transcriptions
Aligned parallel corpora for the following language pair: English-SiSwati. The data is given as four