Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

UniversalDependencies/UD_Northwest_Gbaya-Autogramm

Domain:

natural language processing

Record type:

dataset
Creator:
Uni
Host:
# Summary A Universal Dependencies corpus for Northwest Gbaya, a member of the Gbaya branch of the Atlantic-Congo phylum. The language is mainly spoken by about 250,000 speakers in Central African Republic. # Introduction The treebank is an automatic conversion of the mSUD_Northwest_Gbaya-Autogramm, which was extracted from Paulette Roulon's corpus in Elan format (corpafroas.huma-num.fr). Sentences are annotated with the following metadata: - `sent_id` (which indicates the source file and the segmentation identifier in the source file) - `speaker_id` (which identifies the turn of speech) - `sound_url` (which enables playback of the audio recording) (to be added) - `sent_timecode` (which enables playback of the sentence) - `text` (lexical tokenization) - `text_fr` (French interpretation) # Structure This version of the treebank is a dependency parsing of three files of the original corpus. The original data are spoken data, which were originally segmented in interpausal units, and interlinearized, translated and glossed in Elan. For the syntactic treebank, a re-alignment was done using the illocutionary unit as a sentence. The UD Northwest Gbaya treebank counts 2417 words for 403 sentences. # Reference … # Acknowledgments This treebank was produced as part of the Autogramm ANR project. With special thanks to Christian Chanard for the conversion from Elan, Sylvain Kahane for the mSUD annotation, Aleksandra Miletic and Bruno Guillaume for the conversion from mSUD to UD. ## References * (citation) # Changelog * 2024-11-15 v2.15 * Initial release in Universal Dependencies. === Machine-readable metadata (DO NOT REMOVE!) ================================ Data available since: UD v2.15 License: CC BY-SA 4.0 Includes text: yes Parallel: no Genre: spoken Lemmas: manual native UPOS: manual native XPOS: not available Features: manual native Relations: manual native Contributors: Roulon, Paulette Contributing: elsewhere Contact: …

Visit

github.com

Tasks

dependency parsingparsing

Languages

GbayaGbaya, Northwest

Similar

Autogramm/Gbayasurfacesyntacticud/mSUD_Northwest_Gbaya-AutogrammSUD_Zaar-Autogramm SUD_Zaar-Autogramm: A syntactic treebank of Zaar (aka Saya), a Chadic language of NigeriaUniversalDependencies/UD_Amharic-InkuUniversalDependencies/UD_Nkore-ENkoreTBUniversalDependencies/UD_Malagasy-Hazo

Autogramm/Gbaya

# SUD_Gbaya-Autogramm This repository contains the SUD annotation of the **Gbaya-Autogramm** corpus

surfacesyntacticud/mSUD_Northwest_Gbaya-Autogramm

Treebank for the Gbaya language # Beja-Autogramm This repository contains the SUD annotation of th

SUD_Zaar-Autogramm SUD_Zaar-Autogramm: A syntactic treebank of Zaar (aka Saya), a Chadic language of Nigeria

A Universal Dependencies corpus for Zaar (aka Sayanci), a member of the Chadic branch

UniversalDependencies/UD_Amharic-Inku

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...

UniversalDependencies/UD_Nkore-ENkoreTB

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...

UniversalDependencies/UD_Malagasy-Hazo

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...