Logo Lanfrica

Morphologically annotated corpus for Sesotho

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Gaustad, Tanja
Éditeur:
McKellar, Cindy
Éditeur:
Centre for Text Technology (CTexT)
Hôte:avatar
NCHLT corpus of morphologically annotated tokens in Sesotho converted to the tags used during phases 1 and 2 of the SADiLaR-II project. The data is given as txt files. Each line consists of a token and the corresponding morphological analysis, tab separated. The file for Sesotho contains a total of 73,727 tokens. All the data has been automatically converted, then manually checked and re-annotated where necessary by linguistic experts as well as quality controlled. Please see the included protocol for more details on the morphological tags used.