Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kekem-ALCAM-MultimodalDataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
Kekem-ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Kekem variety of Mbo (ISO 639-3: mbo), a Bantu language spoken in and around the Kekem subdivision of the Haut-Nkam Department of Cameroon's West Region. Mbo, and the Kekem dialect in particular, remains virtually absent from existing grammatical descriptions, computational resources, and lexicographical tools. The dataset comprises three closely aligned components: (i) a structured datasheet containing carefully selected example sentences and lexical entries reflecting attested usage in the Kekem variety of Mbo; (ii) high-quality audio recordings of these entries and sentences, produced by a native speaker across four recording sessions; and (iii) per-session audio–sentence mapping files enabling precise alignment between the textual and acoustic data. The dataset's primary added value lies in its explicit focus on the Kekem variety of Mbo, a language that, like many other Bantu languages of Cameroon situated at the crossroads of linguistic regions, remains essentially absent from reference grammars, dictionaries, educational materials, and language technology resources. Although Kekem is administratively located in the West Region, the indigenous Mbo people of this area do not identify as Bamiléké; rather, they share closer linguistic and historical ties with the Sawa (coastal) peoples. The Kekem dialect displays a range of phonological and morphosyntactic features characteristic of Mbo, including a complex system of vowel contrasts, tonal distinctions, prenasalised consonants and glottal closure markers, all of which are essential for understanding the language's structural specificity and are rarely documented in machine-readable form. In this sense, the dataset contributes to a more inclusive and granular representation of African linguistic diversity. From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The parallel availability of text in Kekem (in IPA transcription) and in French, alongside aligned speech, makes the dataset suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment, pronunciation modelling, and multimodal language learning tools. At the same time, the structured datasheet supports linguistic analysis, comparison with other Mbo dialects and related coastal Bantu languages, and pedagogical uses in teacher training and language revitalisation contexts. The dataset was collected through the Atlas Linguistique du Cameroun (ALCAM) questionnaire framework, designed to gather basic lexical and grammatical information about Cameroonian national languages. The audio recordings were produced at the École Normale Supérieure de Yaoundé (ENS-Yaoundé) in June 2026, in the framework of the Mozilla Data Collective project. More broadly, the Kekem-ALCAM-MultimodalDataset exemplifies an approach to African language resources that highlights fluidity, orality, phonological richness, and community-based linguistic practice.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processing

Languages

MboMbo

Tags

mdcmozilla data collectiveNLPMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Baka-ALCAM-MultimodalDatasetAkoose-ALCAM-MultimodalDatasetBasaa-ALCAM-MultimodalDatasetDiboum-ALCAM-MultimodalDatasetNgemba-ALCAM-MultimodalDatasetBatanga-ALCAM-MultimodalDataset

Baka-ALCAM-MultimodalDataset

Baka-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and t

Akoose-ALCAM-MultimodalDataset

This dataset comprises a datasheet of Akoose (bss) lexical entries collected from the 'Western-Bakos

Basaa-ALCAM-MultimodalDataset

This dataset comprises a datasheet of lexical entries in Basaa, accompanied by illustrative sentence

Diboum-ALCAM-MultimodalDataset

Diboum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Ngemba-ALCAM-MultimodalDataset

Ngemba_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Batanga-ALCAM-MultimodalDataset

Batanga-ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the