Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

UniversalDependencies/UD_Naija-NSC

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Uni
Hôte:
Naija (Nigerian Pidgin) data. # SUD_Naija-NSC # Summary A Universal Dependencies corpus for spoken Naija (Nigerian Pidgin). # Introduction The corpus is based on dialogues and monologues and comprises 9,242 sentences and 140,729 tokens. Sentences are annotated with the following metadata : + sent_id (which also indicates the sample file) + text + text_en (English translation) + text_ortho (A simplified version of text where macrosyntactic annotation has been replaced by standard punctuation) + speaker_id (from the NaijaSynCor Metadata) + sound_url (links to the corresponding sound file, AlignBegin and AlignEnd features give the miliseconds that allow for a positioning in the soundfile) # Structure The text has been transcribed mostly following English spelling conventions for lexical words. Grammatical words have been transcribed following consensual conventions elaborated by the annotators. The text is segmented into illocutionary units. The end of illocutionary units is indicated by a double slash (//). The sentence nucleus containing the predicate is separated from dislocated units by "lesser than" signs ( ) from right-dislocated units. Paradigmatic lists (coordinations, appositions, and disfluencies) are marked with curly breackets, each conjunct being separated by the pipe symbol (|). Further details can be found on the "Macrosyntactic annotation guide". The treebank is developed in SUD (surfacesyntacticud.github.io) and is converted automatically into UD. # Deviations from UD - We distinguish arguments and modifiers in the `obl` relation: `obl:arg`, `obl:mod` - We use `compound:redup` for reduplications and `compound:svc` for serial verb constructions. - We distinguish 6 types of parataxis: `parataxis:conj`, `parataxis:discourse`, `parataxis:dislocated`, `parataxis:insert`, `parataxis:obj`, `parataxis:parenth`. Consult the language specific documentation for further details concerning subtypes. # Acknowledgments The treebank was created within the NaijaSynCor projec …

Visit

github.com

Tasks

dependency parsingparsing

Similaires

UniversalDependencies/UD_Hausa-EasternAutogrammUniversalDependencies/UD_Amharic-SAMTAUniversalDependencies/UD_Ruuli-RDTUniversalDependencies/UD_Kabyle-ADPTUniversalDependencies/UD_Khoekhoe-KDTUniversalDependencies/UD_Hausa-SouthernAutogramm

UniversalDependencies/UD_Hausa-EasternAutogramm

# Summary This treebank contains data of the Autogramm project, for the (Kano) Eastern dialect of H

UniversalDependencies/UD_Amharic-SAMTA

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...

UniversalDependencies/UD_Ruuli-RDT

# Summary UD_Ruuli-RDT is a Universal Dependencies (UD) treebank for the Ruruuli-Lunyala (Ruuli) la

UniversalDependencies/UD_Kabyle-ADPT

# Summary ... 1-2 sentences (see release checklist for README guidelines) ... # Introduction ...

UniversalDependencies/UD_Khoekhoe-KDT

# Summary UD\_Khoekhoe-KDT is a Universal Dependencies (UD) treebank for the Khoekhoegowab (Khoekho

UniversalDependencies/UD_Hausa-SouthernAutogramm

# Summary This treebank contains data of Southern Autogramm, for the Zaria dialect of Nigeria (Sout