Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Babute-Njore_ALCAM-MultimodalDataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
Babute-Njore_ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and technological enhancement of Babouté (also known as Vute, Bute, Wute, Voute; ISO 639-3: vut), a Bantoid language spoken by farming communities across the Centre, Adamawa and East Regions of Cameroon. This release documents the variety of Babouté spoken in Njoré, a Vute-speaking locality in the Mbandjock area (Haute-Sanaga division, Centre Region). Babouté remains sparsely represented in computational language resources despite an estimated 21,000 speakers and a documented standard orthography. The dataset comprises three closely aligned components: (i) a datasheet containing lexical entries and example sentences reflecting attested usage in Babouté as spoken in Njoré; (ii) high-quality audio recordings of a substantial subset of these entries, produced by a native speaker; and (iii) explicit audio-sentence mapping files enabling precise alignment between the textual and acoustic data. The dataset's primary added value lies in its explicit documentation of the Njoré variety of Babouté, a Bantoid language of central Cameroon that — despite a comparatively larger speaker base and an existing General Alphabet of Cameroonian Languages (GACEL)-based orthography (standardized in 1979) — remains, like many regional languages of Cameroon, poorly represented in modern computational and pedagogical resources at the level of specific local varieties. The parallel availability of text in Babouté and in French, together with aligned speech for a substantial subset of entries, makes the dataset suitable for a range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment and pronunciation modelling. From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The datasheet's word-for-word parsing of both the Babouté and French example sentences further supports morphological analysis and glossed-corpus studies. At the same time, the structured datasheet supports basic lexicographic and grammatical documentation, and pedagogical uses in teacher training and language revitalisation contexts. More broadly, the Babute-Njore_ALCAM-MultimodalDataset exemplifies an approach to Cameroonian language resources that documents linguistic variation at the level of specific villages and speech communities, rather than only at the level of the named language as a whole.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

TukiVute

Tags

mdcmozilla data collectiveNLPMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Ewondo_Fong_ALCAM-MultimodalDatasetYezoum_ALCAM-MultimodalDatasetBulu_ALCAM-MultimodalDatasetMvele_ALCAM-MultimodalDatasetGhomala-MultimodalDatasetBaka-ALCAM-MultimodalDataset

Ewondo_Fong_ALCAM-MultimodalDataset

Ewondo_Fong_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to

Yezoum_ALCAM-MultimodalDataset

Yezoum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Bulu_ALCAM-MultimodalDataset

ALCAM-Bulu-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the doc

Mvele_ALCAM-MultimodalDataset

Mvele_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the do

Ghomala-MultimodalDataset

Ghomala-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the docume

Baka-ALCAM-MultimodalDataset

Baka-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and t