Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.
The Afrikaans Ner Corpus is an Afrikaans dataset developed by The Centre for Text Technology (CTexT)
African Speech Technology speech and transcription data for the Afrikaans-Afrikaans database. The "
This is the Afrikaans language part of the NCHLT Speech Corpus of the South African languages. Langu
Orthographic and phonemically aligned transcriptions
Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems.