Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Yoruba-English Code-Switching (YECS) Corpus | Mozilla Data Collective

Domaine:

natural language processing

Type de record:

dataset
The Yoruba-English Code-Switching (YECS) Corpus is a comprehensive, ~120-hour dataset designed to capture the natural linguistic phenomenon of intra-sentential code-mixing. Curated by the LynguaTech Innovative Foundation (LyngualLabs), this dataset provides nearly 100,000 validated audio-text pairs recorded by 140 demographically diverse bilingual speakers in Nigeria. It features clean speech recordings paired with full Yoruba orthography (including verified tonal marks and diacritics), word-level language identification tags, and rich metadata spanning 16 semantic domains and 7 emotion categories. The dataset is explicitly partitioned to prevent data contamination, serving as a highly stratified, robust benchmark for low-resource speech technologies.

Visit

mozilladatacollective.com

Tasks

code switching

Languages

Yoruba

Tags

mdcmozillaLyngualLabs

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Yoruba/english conversational code‐switching as a conversational strategyVariation and language engineering in Yoruba-English code-switchingCorpus of English and Nigerian Pidgin Code-switching (CENCOS)AremuAdeolaJr/An-NLP-Deep-Dive-into-Yoruba-English-Code-SwitchingCode‐Switching: Amharic‐EnglishThe Manifestation of Code Switching Among 3 Year-Old Yoruba/English Semilinguals

Yoruba/english conversational code‐switching as a conversational strategy

Variation and language engineering in Yoruba-English code-switching

This study deals with the identification and characterization of the variable features of code-switc

Corpus of English and Nigerian Pidgin Code-switching (CENCOS)

This dataset was compiled from a fieldwork in Nigeria in 2019. It features naturally occurring spoke

AremuAdeolaJr/An-NLP-Deep-Dive-into-Yoruba-English-Code-Switching

# An NLP Deep Dive into Yoruba-English Code-Switching: Synthetic Corpus Construction, Human Quality

Code‐Switching: Amharic‐English

The Manifestation of Code Switching Among 3 Year-Old Yoruba/English Semilinguals