Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Bulu_ALCAM-MultimodalDataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
ALCAM-Bulu-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Bulu language, a Bantu language spoken in southern Cameroon. The dataset comprises three closely aligned components: (i) a structured datasheet containing carefully selected example sentences and lexical entries reflecting usage in the Bulu language; (ii) high-quality audio recordings of these sentences and lexical items, produced by a native speaker; and (iii) an explicit audio–sentence mapping file enabling precise alignment between the textual and acoustic data. The dataset's primary added value lies in its explicit focus on Bulu, a language which, despite having a significant speaker community and a rich literary tradition, remains severely under-represented in digital language resources and speech technology infrastructures. The dataset captures the phonological, morphological and lexical properties of Bulu through a structured elicitation methodology, and is designed to serve as a foundational resource for both linguistic analysis and the development of speech technology applications. From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The parallel availability of text in Bulu alongside aligned speech makes the dataset suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), forced alignment, pronunciation modelling, and multimodal language learning tools. The structured datasheet further supports linguistic analysis, morphological study, contrastive studies with related varieties of the Beti-Fang group, and pedagogical uses in teacher training and language revitalisation contexts. More broadly, the ALCAM-Bulu-MultimodalDataset exemplifies an approach to African language resources that highlights socially embedded linguistic practice, phonological precision, and community-based documentation.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

Bulu

Tags

mdcmozilla data collectiveNLPMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Ewondo_Fong_ALCAM-MultimodalDatasetYezoum_ALCAM-MultimodalDatasetMvele_ALCAM-MultimodalDatasetGhomala-MultimodalDatasetBaka-ALCAM-MultimodalDatasetAkoose-ALCAM-MultimodalDataset

Ewondo_Fong_ALCAM-MultimodalDataset

Ewondo_Fong_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to

Yezoum_ALCAM-MultimodalDataset

Yezoum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Mvele_ALCAM-MultimodalDataset

Mvele_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the do

Ghomala-MultimodalDataset

Ghomala-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the docume

Baka-ALCAM-MultimodalDataset

Baka-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and t

Akoose-ALCAM-MultimodalDataset

This dataset comprises a datasheet of Akoose (bss) lexical entries collected from the 'Western-Bakos