Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Gbaya-Lay_ALCAM-MultimodalDataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
Gbaya-Lay_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of Gbaya-Lay (Làì), a variety of Northwest Gbaya (ISO 639-3: gya), a Niger-Congo language spoken in Cameroon and the Central African Republic. Gbaya-Lay is a localised and socially embedded speech form that is rarely represented in standard grammatical descriptions or lexicographical resources. The dataset comprises three closely aligned components: (i) a structured datasheet containing carefully selected example sentences reflecting usage in Gbaya-Lay; (ii) high-quality audio recordings of these sentences, produced by a native speaker across three recording sessions; and (iii) explicit audio–sentence mapping files enabling precise alignment between the textual and acoustic data. The dataset's primary added value lies in its explicit focus on the Gbaya-Lay variety of Northwest Gbaya. Gbaya-Lay (Làì) is classified as a dialect of Northwest Gbaya (gya) by the Ethnologue (ethnologue.com) and is referenced in the standard atlases of Cameroon's languages: the Atlas Linguistique du Cameroun by Breton and Bikia Fohtung (1991) and the Atlas Linguistique de l'Afrique Centrale: le Cameroun by Bibam Bikoi (2012). Gbaya-Lay is geographically restricted to a small area north of Mbodomo, in Cameroon, distinguishing it from the more widely spoken Gbaya-Kara (Kàrà) variety. Like many geographically and socially situated language varieties, Gbaya-Lay typically remains invisible in reference grammars, dictionaries and educational materials that often privilege better-documented or more standardised forms of the language. The dataset captures micro-variation in phonetics, phonology, morphosyntax and lexical choice that are essential for understanding socially situated linguistic practices rather than a homogeneous, abstract system. In this sense, the dataset contributes to a more inclusive representation of linguistic diversity within the Northwest Gbaya speech community. From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The parallel availability of text in Gbaya-Lay and in French, alongside aligned speech, makes the dataset suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment, pronunciation modelling and multimodal language learning tools. At the same time, the structured datasheet supports linguistic analysis, contrastive studies with other language varieties and pedagogical uses in teacher training and language revitalisation contexts. More broadly, the Gbaya-Lay_ALCAM-MultimodalDataset exemplifies an approach to African language resources that highlights fluidity, longitudinal variation, orality and community-based practice.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionmachine translationspeech processingtext to speech

Languages

GbayaGbaya-MbodomoGbaya, NorthwestGbaya, Southwest

Tags

mdcmozilla data collectiveNLPMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Ewondo_Fong_ALCAM-MultimodalDatasetYezoum_ALCAM-MultimodalDatasetBulu_ALCAM-MultimodalDatasetMvele_ALCAM-MultimodalDatasetGhomala-MultimodalDatasetBaka-ALCAM-MultimodalDataset

Ewondo_Fong_ALCAM-MultimodalDataset

Ewondo_Fong_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to

Yezoum_ALCAM-MultimodalDataset

Yezoum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Bulu_ALCAM-MultimodalDataset

ALCAM-Bulu-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the doc

Mvele_ALCAM-MultimodalDataset

Mvele_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the do

Ghomala-MultimodalDataset

Ghomala-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the docume

Baka-ALCAM-MultimodalDataset

Baka-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and t