Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

MELANZ: A Trilingual Kreol Morisien–English–French Code-Switching SpeechCorpus for Automatic Speech Recognition

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
YarGobPud
Éditeur:
UniProPro
Éditeur:
CCSD
Hôte:avatar

MELANZ is the first code-switching (CS) speech corpus for Kreol Morisien (KM), a French-lexifier Creole language spoken natively by over 90% of the Mauritian population. In practice, KM speakers routinely alternate between KM, English, and French within a single utterance, yet no dedicated speech resource exists to reflect this multilingual reality. The MELANZ corpus comprises approximately 22 hours of broadcast news speech from 76 episodes of the Zournal Kreol television programme. The recordings were segmented into 4,037 utterances totalling 197,028 words, with all transcriptions manually verified. A 2-hour evaluation subset (271 utterances, 16,098 words) was annotated with word-level language identification tags, achieving an inter-annotator agreement of Cohen's kappa = 0.966. Corpus analysis confirms the trilingual nature of KM speech: 91.1% of utterances contain at least one language switch, 46.1% contain words from all three languages simultaneously, and the average Code-Mixing Index is 0.189. Baseline ASR experiments with three architecturally distinct models (Whisper large-v3, MMS, wav2vec2-XLSR-1B) reveal that French words are consistently the most difficult to recognise (WER 31.3-47.9%), error rates at switch points are 130-157% higher than elsewhere, and recognition degrades monotonically with CS intensity. MELANZ fills a critical gap as the first CS speech resource for any Creole language.

Visit

hal.science

Tasks

automatic speech recognitioncode switchinglanguage identificationspeech processing

Languages

MorisyenSeychelles French Creole

Tags

automatic speech recognitionlow-resource languagetrilinguallanguage identificationCCS Concepts:Computing methodologies → Speech recognitionNatural language processing• Information systems → Language models code-switchingKreol Morisienspeech corpus+4

Licenses

https://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/OpenAccess

Similaires

Kreol Morisien to English and English to Kreol Morisien Translation System using Attention and Transformer ModelInvestigations on Speech Recognition Systems for Low-Resource Dialectal Arabic-English Code-Switching SpeechAutomatic Translation Between Kreol Morisien and English Using the Marian Machine Translation FrameworkThe evaluation of a code-switched Sepedi-English automatic speech recognition systemEffects of Language Modelling for Sepedi-English Code-Switched Speech in Automatic Speech Recognition SystemCAFE: Spontaneous code-switching speech dataset in Algerian dialect, French and English

Kreol Morisien to English and English to Kreol Morisien Translation System using Attention and Transformer Model

Investigations on Speech Recognition Systems for Low-Resource Dialectal Arabic-English Code-Switching Speech

Code-switching (CS), defined as the mixing of languages in conversations, has become a worldwide phe

Automatic Translation Between Kreol Morisien and English Using the Marian Machine Translation Framework

Kreol Morisien is a vibrant and expressive language that reflects the multicultural heritage of Maur

The evaluation of a code-switched Sepedi-English automatic speech recognition system

Speech technology is a field that encompasses various techniques and tools used to enable machines t

Effects of Language Modelling for Sepedi-English Code-Switched Speech in Automatic Speech Recognition System

CAFE: Spontaneous code-switching speech dataset in Algerian dialect, French and English