Logo Lanfrica

SUD_Zaar-Autogramm SUD_Zaar-Autogramm: A syntactic treebank of Zaar (aka Saya), a Chadic language of Nigeria

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
CarKahGui
Éditeur:
LanModSemANR
Éditeur:
CCSD
Hôte:avatar
A Universal Dependencies corpus for Zaar (aka Sayanci), a member of the Chadic branch of the Afro-Asiatic phylum. The language is mainly spoken by about 200,000 speakers in the Bogoro and Tafawa Balewa local governments of Bauchi State, Nigeria. This version of the treebank is a dependency parsing of the original corpus first three files.The original data are spoken data, which were originally segmented in interpausal units, and interlinearized, translated and glossed in Elan. For the syntactic treebank, a re-alignment was done using the illocutionary unit as a sentence. Tokens comprize words and affixes (preceded by a "=" sign) when those bear a syntactic function. Punctuation tokens (e.g. <, >, //, etc.) organise the illocutionary unit into: pre-nucleus < nucleus > post-nucleus //.The UD Zaar treebank counts 8441 tokens for 817 sentences.