Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Ghomala-MultimodalDataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
Ghomala-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the ɣɔmáláʔ language (Ghɔmala'), as documented in the Bandjoun and Bamougoum villages of the West Region of Cameroon. ɣɔmáláʔ is a Grassfields Bantu language of the Mbam-Nkam branch of the Bantoid family. It is rarely represented in existing computational resources. The dataset was compiled in the context of doctoral research on the forms and functions of ritual language in the ɣɔmáláʔ-speaking community (2020-2021). The dataset comprises three closely aligned components: (i) a structured fieldwork datasheet containing 376 IPA-transcribed example sentences extracted from recorded ritual speech events, together with their word-for-word parsing, interlinear glosses and French translations; (ii) 369 high-quality audio recordings of these sentences, produced by a native speaker of ɣɔmáláʔ across four recording sessions; and (iii) per-session audio–sentence mapping files enabling precise alignment between the textual and acoustic data. The dataset additionally includes a bilingual parallel corpus (Ghomala–French) in TSV format, derived from the same source material. The ritual texts captured in this dataset originate from five distinct ceremonial contexts documented in the Bandjoun and Bamougoum speech communities: rites of intercession for healing, goat sacrifice rituals, dowry ceremonies, purification rites, and installation rites. This breadth of ritual registers makes the dataset particularly valuable for studying specialised and formulaic language use in a tonal Grassfields Bantu language. From a methodological perspective, the dataset bridges language documentation and language technology. The parallel availability of IPA-transcribed text in ɣɔmáláʔ and French, alongside aligned speech, makes it suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment, pronunciation modelling and multimodal language learning tools. The structured datasheet, with its interlinear glosses and word-level parsing, additionally supports linguistic analysis, contrastive studies with other Grassfields Bantu varieties, and pedagogical uses in teacher training and language revitalisation contexts. The phonological inventory documented in this dataset — including a complex tonal system, ejective consonants, nasal vowels and vowel harmony — reflects the full structural richness of ɣɔmáláʔ, and contributes to a more inclusive and granular representation of African linguistic diversity in language technology resources.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionmachine translationspeech processingtext to speech

Languages

Ghomálá’

Tags

mdcmozilla data collectiveNLPTSVMP3

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Ewondo_Fong_ALCAM-MultimodalDatasetYezoum_ALCAM-MultimodalDatasetBulu_ALCAM-MultimodalDatasetMvele_ALCAM-MultimodalDatasetFrench to Ghomala (Bandjoun)Sample Ghomala-TTS-Dataset

Ewondo_Fong_ALCAM-MultimodalDataset

Ewondo_Fong_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to

Yezoum_ALCAM-MultimodalDataset

Yezoum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the d

Bulu_ALCAM-MultimodalDataset

ALCAM-Bulu-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the doc

Mvele_ALCAM-MultimodalDataset

Mvele_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the do

French to Ghomala (Bandjoun)

Dataset containing translation of a given set of French words and expressions to Ghomala, the native

Sample Ghomala-TTS-Dataset

Sample-Ghomala-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technolog