Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

UniversalDependencies/UD_Ruuli-RDT

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Uni
Hôte:
# Summary UD_Ruuli-RDT is a Universal Dependencies (UD) treebank for the Ruruuli-Lunyala (Ruuli) language. The annotation was converted from interlinear glossed text and manually annotated for syntactic relations. The treebank includes texts from various sources: conversations, oral folktales, biographic monologue, movie subtitles, grammar examples, and factual prose. The treebank contains approximately 6,000 tokens. # Introduction The UD_Ruuli-RDT treebank consists of texts recorded in Ruuli or translated into it by native speakers, and subsequently glossed and annotated. The included texts are: * spokenBio_Nakasongola1 (795 words): Biographic monologue on childhood years, schooling, work, and other life experiences * spokenConv_Nakasongola1b (1223 words): Conversation between two speakers, a male and a female, about taking care of their elderly parents * spokenConv_Nakasongola2 (1544 words): Conversation between two females about socio-economic issues * spokenTale_Gweero (595 words): A traditional oral folktale about the cow who got in trouble with the lion and the hare who helped the cow * spokenTale_Sokoso (373 words): A traditional oral folktale about a woman who mistreated her mother-in-law * film_Inception (469 words): An excerpt from the translated subtitles for the film *Inception* (2010) * grammar_Syntax (936 words): Language examples from *A dictionary and grammatical sketch of Ruruuli-Lunyala* (Namyalo et al. 2021) * nonfiction_Aniinire (366 words): An excerpt from factual prose on the history and traditions of the language speakers All sentences were converted from interlinear glossed text into CoNLL-U format using a custom conversion script. The syntactic relations were subsequently manually annotated following the UD framework. Sentences from written texts and conversations were shuffled to anonymize the data. # Genre Classification * Spoken (incl. conversations, oral folktales, and biographic monologue): sentence IDs start with `spoken` * Fic …

Visit

github.com

Tasks

dependency parsingparsing

Languages

NyalaRuruuli-Runyala

Similaires

UniversalDependencies/UD_Hausa-EasternAutogrammUniversalDependencies/UD_Amharic-SAMTAUniversalDependencies/UD_Northwest_Gbaya-AutogrammUniversalDependencies/UD_Kabyle-ADPTUniversalDependencies/UD_Khoekhoe-KDTUniversalDependencies/UD_Hausa-NorthernAutogramm

UniversalDependencies/UD_Hausa-EasternAutogramm

# Summary This treebank contains data of the Autogramm project, for the (Kano) Eastern dialect of H

UniversalDependencies/UD_Amharic-SAMTA

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...

UniversalDependencies/UD_Northwest_Gbaya-Autogramm

# Summary A Universal Dependencies corpus for Northwest Gbaya, a member of the Gbaya branch of the

UniversalDependencies/UD_Kabyle-ADPT

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...

UniversalDependencies/UD_Khoekhoe-KDT

# Summary UD\_Khoekhoe-KDT is a Universal Dependencies (UD) treebank for the Khoekhoegowab (Khoekho

UniversalDependencies/UD_Hausa-NorthernAutogramm

# Summary This treebank contains data of Northern Autogramm, for the Ader dialect of Niger Republic