Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Ngemba-ALCAM-MultimodalDataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
Ngemba_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Ngemba language, also referred to in the literature as Ghomala-Ouest (Breton and Bikia Fohtung 1991). Ngemba is a Grassfields Bantu language spoken in the West Region of Cameroon and is rarely represented in existing standard grammatical descriptions, computational resources or lexicographical tools. The dataset comprises three closely aligned components: (i) a structured datasheet containing carefully selected example sentences and lexical entries reflecting attested usage in Ngemba; (ii) high-quality audio recordings of these entries, produced by a native speaker; and (iii) an explicit audio–sentence mapping file enabling precise alignment between the textual and acoustic data. The dataset's primary added value lies in its explicit focus on Ngemba, a language that, like many other Grassfields Bantu languages, remains virtually absent from reference grammars, dictionaries, educational materials and language technology resources. The dataset captures a range of phonological and morphosyntactic features characteristic of Ngemba, including a complex system of vowel harmony, nasal vowels, ejective consonants and lexical tone, all of which are essential for understanding the language's structural specificity and are rarely documented in machine-readable form. In this sense, the dataset contributes to a more inclusive and granular representation of African linguistic diversity. From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The parallel availability of text in Ngemba and in French, alongside aligned speech, makes the dataset suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment, pronunciation modelling and multimodal language learning tools. At the same time, the structured datasheet supports linguistic analysis, contrastive studies with other Grassfields Bantu varieties and pedagogical uses in teacher training and language revitalisation contexts. More broadly, the Ngemba_ALCAM-MultimodalDataset exemplifies an approach to African language resources that highlights fluidity, orality, phonological richness and community-based linguistic practice.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionmachine translationspeech processingtext to speech

Languages

Ngemba

Tags

mdcmozilla data collectiveNLPMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Diboum-ALCAM-MultimodalDatasetBatanga-ALCAM-MultimodalDatasetAkoose-ALCAM-MultimodalDatasetKekem-ALCAM-MultimodalDatasetBasaa-ALCAM-MultimodalDatasetBaka-ALCAM-MultimodalDataset

Diboum-ALCAM-MultimodalDataset

Diboum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Batanga-ALCAM-MultimodalDataset

Batanga-ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the

Akoose-ALCAM-MultimodalDataset

This dataset comprises a datasheet of Akoose (bss) lexical entries collected from the 'Western-Bakos

Kekem-ALCAM-MultimodalDataset

Kekem-ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the do

Basaa-ALCAM-MultimodalDataset

This dataset comprises a datasheet of lexical entries in Basaa, accompanied by illustrative sentence

Baka-ALCAM-MultimodalDataset

Baka-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and t