Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated wi
Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.
Monolingual corpus for Setswana. The data is given as a single UTF-8 text file, with each segment on
Orthographic and phonemically aligned transcriptions
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test sui
Contains training and testing data for Genre Classification for Setswana.