Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Batanga-ALCAM-MultimodalDataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
Batanga-ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Batanga language (ISO 639-3: bnm). Batanga is a Bantu language spoken along the Atlantic coast of the South Region of Cameroon and is rarely represented in existing grammatical descriptions, computational resources or lexicographical tools. The dataset is published in two successive releases: the present release covers the Banoho (banɔɔ) dialect; a companion datasheet for the Bapuku dialect will be integrated in a forthcoming release. The complete dataset will comprise three closely aligned components for each dialect: (i) a structured datasheet containing carefully selected example sentences and lexical entries reflecting attested usage in Batanga; (ii) high-quality audio recordings of these entries, produced by a native speaker; and (iii) an explicit audio–sentence mapping file enabling precise alignment between the textual and acoustic data. The dataset's primary added value lies in its explicit focus on Batanga, a language that, like many other coastal Bantu languages of Cameroon, remains virtually absent from reference grammars, dictionaries, educational materials and language technology resources. The Banoho and Bapuku varieties display a range of phonological and morphosyntactic features characteristic of the Cameroonian coastal Bantu area, including a complex system of vowel contrasts, nasal vowels, and lexical tone, all of which are essential for understanding the language's structural specificity and are rarely documented in machine-readable form. In this sense, the dataset contributes to a more inclusive and granular representation of African linguistic diversity. From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The parallel availability of text in Batanga and in French, alongside aligned speech, makes the dataset suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment, pronunciation modelling and multimodal language learning tools. At the same time, the structured datasheet supports linguistic analysis, contrastive studies between the Banoho and Bapuku varieties, comparison with related coastal Bantu languages, and pedagogical uses in teacher training and language revitalisation contexts. More broadly, the Batanga-ALCAM-MultimodalDataset exemplifies an approach to African language resources that highlights fluidity, orality, phonological richness and community-based linguistic practice.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

Batanga

Tags

mdcmozilla data collectiveNLPMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Baka-ALCAM-MultimodalDatasetAkoose-ALCAM-MultimodalDatasetBasaa-ALCAM-MultimodalDatasetDiboum-ALCAM-MultimodalDatasetNgemba-ALCAM-MultimodalDatasetKekem-ALCAM-MultimodalDataset

Baka-ALCAM-MultimodalDataset

Baka-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and t

Akoose-ALCAM-MultimodalDataset

This dataset comprises a datasheet of Akoose (bss) lexical entries collected from the 'Western-Bakos

Basaa-ALCAM-MultimodalDataset

This dataset comprises a datasheet of lexical entries in Basaa, accompanied by illustrative sentence

Diboum-ALCAM-MultimodalDataset

Diboum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Ngemba-ALCAM-MultimodalDataset

Ngemba_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Kekem-ALCAM-MultimodalDataset

Kekem-ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the do