Badja-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and technological enhancement of Badja (also spelled Badjia; dialect name Bakjo), one of the named dialects of Ewondo (also known as Beti, Yaunde; ISO 639-3: ewo), a Bantu language spoken primarily by the Beti people across Cameroon's Centre Region. Ewondo remains one of Cameroon's better-documented national languages overall, with an estimated 577,700 speakers (1982 figure, the most recent census-based estimate) and a role as a widely used trade language around Yaoundé. Despite this, its constituent dialects — including Badja — remain, like many locally defined varieties of Cameroon's national languages, sparsely represented in computational language resources at the level of the specific dialect rather than the language as a whole. The dataset comprises three closely aligned components: (i) a datasheet containing lexical entries and example sentences reflecting attested usage in Badja; (ii) audio recordings of a substantial subset of these entries, produced by a native speaker; and (iii) explicit audio-sentence mapping files enabling precise alignment between the textual and acoustic data.
The dataset's primary added value lies in its explicit documentation of Badja (Badjia), a dialect of Ewondo that — despite Ewondo's comparatively large speaker base, its status as a major Beti-Fang trade language, and the existence of a harmonized practical orthography developed for national-language teaching — remains, at the level of this specific dialect, poorly represented in modern computational and pedagogical resources. The parallel availability of text in Badja and in French, together with aligned speech for a subset of entries, makes the dataset suitable for a range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment and pronunciation modelling.
From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The datasheet's word-for-word parsing of both the Badja and French example sentences further supports morphological analysis and glossed-corpus studies. At the same time, the structured datasheet supports basic lexicographic and grammatical documentation, and pedagogical uses in teacher training and language revitalisation contexts.
More broadly, the Badja-ALCAM-MultimodalDataset exemplifies an approach to Cameroonian language resources that documents linguistic variation at the level of specific dialects and speech communities, rather than only at the level of the named language (Ewondo) as a whole.