Logo Lanfrica

NCHLT Sepedi Phrase Chunk Annotated Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
D.J. PrinslooRoald Eiselen
Éditeur:
North-West UniversityCentre for Text Technology (CTexT)
Hôte:avatar
Phrase chunk annotated data for the NCHLT Text Resource Development: Phase II Project. The phrase chunk annotated data is a subset of the 50,000 tokens annotated during the NCHLT text resource development project and consists of a minimum of 15,000 tokens annotated as one of the six phrase types described in the protocol.