Baka-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Baka language (ISO 639-3: bkc). Baka is an Ubanguian (Ubangi) language spoken by forest-based, hunter-gatherer communities in the southeastern regions of Cameroon, and it remains largely absent from computational resources and language technology tools despite its status as a vigorous, actively transmitted language. The dataset comprises three closely aligned components: (i) a datasheet containing lexical entries and example sentences reflecting attested usage in Baka; (ii) high-quality audio recordings of these entries, produced by a native speaker; and (iii) explicit audio-sentence mapping files enabling precise alignment between the textual and acoustic data.
The dataset's primary added value lies in its explicit focus on Baka, a language that, like many other minority and indigenous languages of Cameroon, remains virtually absent from reference grammars, dictionaries, educational materials and language technology resources. As a language traditionally associated with a marginalized forest-based community, Baka is also of particular sociolinguistic interest: an estimated 30% of its vocabulary is not of Ubanguian origin, reflecting extensive borrowing tied to a specialized forest economy (edible and medicinal plants, honey collecting, hunting), alongside sustained contact with neighboring Bantu languages of the region.
From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The parallel availability of text in Baka and in French, alongside aligned speech for a substantial subset of entries, makes the dataset suitable for a range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment and pronunciation modelling. The datasheet's word-for-word parsing of both the Baka and French example sentences further supports morphological analysis and glossed-corpus studies. At the same time, the structured datasheet supports basic lexicographic and grammatical documentation, and pedagogical uses in teacher training and language revitalisation contexts.
More broadly, the Baka-ALCAM-MultimodalDataset exemplifies an approach to African language resources that highlights fluidity, orality, and community-based linguistic practice among an under-documented indigenous community.