Abstract This article employs a 4-million-word diachronic corpus to examine how the expression of
Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.
Monolingual corpus for isiXhosa. The data is given as a single UTF-8 text file, with each segment on
The Corpus of Spoken isiXhosa The Corpus of Spoken isiXhosa consists of transcribed and annotated
This document lists the abbreviations used in the morphological annotation and part of speech taggin