Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Baka-ALCAM-MultimodalDataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
Baka-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Baka language (ISO 639-3: bkc). Baka is an Ubanguian (Ubangi) language spoken by forest-based, hunter-gatherer communities in the southeastern regions of Cameroon, and it remains largely absent from computational resources and language technology tools despite its status as a vigorous, actively transmitted language. The dataset comprises three closely aligned components: (i) a datasheet containing lexical entries and example sentences reflecting attested usage in Baka; (ii) high-quality audio recordings of these entries, produced by a native speaker; and (iii) explicit audio-sentence mapping files enabling precise alignment between the textual and acoustic data. The dataset's primary added value lies in its explicit focus on Baka, a language that, like many other minority and indigenous languages of Cameroon, remains virtually absent from reference grammars, dictionaries, educational materials and language technology resources. As a language traditionally associated with a marginalized forest-based community, Baka is also of particular sociolinguistic interest: an estimated 30% of its vocabulary is not of Ubanguian origin, reflecting extensive borrowing tied to a specialized forest economy (edible and medicinal plants, honey collecting, hunting), alongside sustained contact with neighboring Bantu languages of the region. From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The parallel availability of text in Baka and in French, alongside aligned speech for a substantial subset of entries, makes the dataset suitable for a range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment and pronunciation modelling. The datasheet's word-for-word parsing of both the Baka and French example sentences further supports morphological analysis and glossed-corpus studies. At the same time, the structured datasheet supports basic lexicographic and grammatical documentation, and pedagogical uses in teacher training and language revitalisation contexts. More broadly, the Baka-ALCAM-MultimodalDataset exemplifies an approach to African language resources that highlights fluidity, orality, and community-based linguistic practice among an under-documented indigenous community.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

BakaBaka

Tags

mdcmozilla data collectiveNLPWAVTSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Akoose-ALCAM-MultimodalDatasetBasaa-ALCAM-MultimodalDatasetDiboum-ALCAM-MultimodalDatasetNgemba-ALCAM-MultimodalDatasetBatanga-ALCAM-MultimodalDatasetKekem-ALCAM-MultimodalDataset

Akoose-ALCAM-MultimodalDataset

This dataset comprises a datasheet of Akoose (bss) lexical entries collected from the 'Western-Bakos

Basaa-ALCAM-MultimodalDataset

This dataset comprises a datasheet of lexical entries in Basaa, accompanied by illustrative sentence

Diboum-ALCAM-MultimodalDataset

Diboum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Ngemba-ALCAM-MultimodalDataset

Ngemba_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Batanga-ALCAM-MultimodalDataset

Batanga-ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the

Kekem-ALCAM-MultimodalDataset

Kekem-ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the do