Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CAFE A Novel Code switching Dataset for Algerian Dialect French and English

Domain:

natural language processing

Record type:

paperdataset
Creator:
LacAbbOukKhe
Host:avatar
The paper introduces and publicly releases (Data download link available after acceptance) CAFE -- the first Code-switching dataset between Algerian dialect, French, and english languages. The CAFE speech data is unique for (a) its spontaneous speaking style in vivo human-human conversation capturing phenomena like code-switching and overlapping speech, (b) addresses distinct linguistic challenges in North African Arabic dialect; (c) the CAFE captures dialectal variations from various parts of Algeria within different sociolinguistic contexts. CAFE data contains approximately 37 hours of speech, with a subset, CAFE-small, of 2 hours and 36 minutes released with manual human annotation including speech segmentation, transcription, explicit annotation of code-switching points, overlapping speech, and other events such as noises, and laughter among others. The rest approximately 34.58 hours contain pseudo label transcriptions. In addition to the data release, the paper also highlighted the challenges of using state-of-the-art Automatic Speech Recognition (ASR) models such as Whisper large-v2,3 and PromptingWhisper to handle such content. Following, we benchmark CAFE data with the aforementioned Whisper models and show how well-designed data processing pipelines and advanced decoding techniques can improve the ASR performance in terms of Mixed Error Rate (MER) of 0.310, Character Error Rate (CER) of 0.329 and Word Error Rate (WER) of 0.538. 24 pages, submitted to tallip

Visit

arxiv.org

Tasks

automatic speech recognitioncode switchingspeech processing

Languages

Arabic, Algerian Spoken

Tags

SoundComputation and LanguageAudio and Speech Processing

Similar

CAFE: Spontaneous code-switching speech dataset in Algerian dialect, French and EnglishCAFE: Algerian Arabic, French, and English Code-Switched Conversational Speech DatasetSexism detection: The first corpus in Algerian dialect with a code-switching in Arabic/ French and English

CAFE: Spontaneous code-switching speech dataset in Algerian dialect, French and English

CAFE: Algerian Arabic, French, and English Code-Switched Conversational Speech Dataset

The CAFE dataset (Code-switching Algerian French English) addresses the lack of publicly available r

Sexism detection: The first corpus in Algerian dialect with a code-switching in Arabic/ French and English

In this paper, an approach for hate speech detection against women in Arabic community on social med