Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Yoruba-English Code-Switching (YECS) Corpus | Mozilla Data Collective

Domain:

natural language processing

Record type:

dataset
The Yoruba-English Code-Switching (YECS) Corpus is a comprehensive, ~120-hour dataset designed to capture the natural linguistic phenomenon of intra-sentential code-mixing. Curated by the LynguaTech Innovative Foundation (LyngualLabs), this dataset provides nearly 100,000 validated audio-text pairs recorded by 140 demographically diverse bilingual speakers in Nigeria. It features clean speech recordings paired with full Yoruba orthography (including verified tonal marks and diacritics), word-level language identification tags, and rich metadata spanning 16 semantic domains and 7 emotion categories. The dataset is explicitly partitioned to prevent data contamination, serving as a highly stratified, robust benchmark for low-resource speech technologies.

Visit

mozilladatacollective.com

Tasks

code switching

Languages

Yoruba

Tags

mdcmozillaLyngualLabs

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)