A section of the Pretoria Sepedi Corpus for POS, manually checked for POS tags.
This deliverable contains part-of-speech tagged data from five different text types for Sepedi. The
NCHLT corpora with tokens lemmatised and converted to POS tags used during the SADiLaR-II project fo
This corpus contains POS annotated data in 5 different genres for Sepedi. The text types included
This dataset provides a comprehensive resource for studying Assamese-English code-mixed language, ma
Collection of texts for general linguistic research, in particular for lexicography