Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

UniversalDependencies/UD_Khoekhoe-KDT

Domain:

natural language processing

Record type:

dataset
Creator:
Uni
Host:
# Summary UD\_Khoekhoe-KDT is a Universal Dependencies (UD) treebank for the Khoekhoegowab (Khoekhoe) language. The annotation was performed manually based on glosses. This treebank includes texts from various sources: fiction, grammar, and spoken conversation. The treebank contains **27k tokens**, distributed as follows: - **Training set**: 15k tokens - **Development set**: 2k tokens - **Test set**: 10k tokens # Introduction The UD\_Khoekhoe-KDT treebank consists of various texts translated into Khoekhoe by native speakers, then glossed and annotated. The included texts are: - **grammar\_Cairo**: 20 examples from the Cairo Cicling Corpus) - **grammar\_BivalTyp**: BivalTyp dataset. - **film\_Bridge**: Subtitles from *Bridge of Spies* (2015). - **film\_Titanic**: A section of subtitles from *Titanic* (1997). - **book\_Khomai**: Chapters from the school book *Khomai* (1971), with shuffled sentences. - **conversation\_Windhoek5**: A recorded conversation between two friends, anonymized by shuffling sentences. ## Genre Classification - **Fiction**: Sentence IDs start with *film/book*. - **Grammar**: Sentence IDs start with *grammar*. - **Spoken**: Sentence IDs start with *conversation*. ## Data Splits - **Training set**: Full **grammar\_Cairo**, sections of **grammar\_BivalTyp**, **film\_Bridge**, **book\_Khomai**, **conversation\_Windhoek5** - **Development set**: Sections of **grammar\_BivalTyp**, **film\_Bridge**, **conversation\_Windhoek5** - **Test set**: Full **film\_Titanic**, sections of **grammar\_BivalTyp**, **book\_Khomai**, **conversation\_Windhoek5** # Acknowledgments * This work was supported by the project "Event Packaging in Language" (University of Zurich Global Strategy and Partnerships Funding Scheme, 2023–2026). * We also acknowledge the contributions of members of the broader Khoekhoe documentation project, whose work on data collection, translation, and linguistic analysis made this treebank possible: Roswitha Gases, Mauricius Gariseb, …

Visit

github.com

Tasks

dependency parsingparsing

Languages

Khoekhoe

Similar

UniversalDependencies/UD_Hausa-EasternAutogrammUniversalDependencies/UD_Amharic-SAMTAUniversalDependencies/UD_Northwest_Gbaya-AutogrammUniversalDependencies/UD_Kabyle-ADPTUniversalDependencies/UD_Tunisian_Arabic-NAxLATUniversalDependencies/UD_Hausa-NorthernAutogramm

UniversalDependencies/UD_Hausa-EasternAutogramm

# Summary This treebank contains data of the Autogramm project, for the (Kano) Eastern dialect of H

UniversalDependencies/UD_Amharic-SAMTA

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...

UniversalDependencies/UD_Northwest_Gbaya-Autogramm

# Summary A Universal Dependencies corpus for Northwest Gbaya, a member of the Gbaya branch of the

UniversalDependencies/UD_Kabyle-ADPT

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...

UniversalDependencies/UD_Tunisian_Arabic-NAxLAT

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...

UniversalDependencies/UD_Hausa-NorthernAutogramm

# Summary This treebank contains data of Northern Autogramm, for the Ader dialect of Niger Republic