Logo Lanfrica

CFPR accent metadata

Domain:

natural language processing

Record type:

dataset
Creator:
FABPorAva
Publisher:
Zenodo
Host:avatar

To address the current lack of dataset of diverse regional variations of French, we have extended the Corpus du Français Parlé de nos Régions (CFPR) with explicit accent labels. This extension provides a new metadata to develop and assess speech technology and study phonetic diversity of the French-speaking world.
CFPR is an open-access resource, developed at Sorbonne Université (Avanzi et al., 2019), designed to study regional and social linguistic variation. It consists of audio recordings from 186 speakers, each providing a single recording. The original dataset is rich in metadata, documenting: Demographics (Gender and year of birth), Geography (Country/region of origin, birthplace, and current residence) and other context (Recording date, location, and the speaker’s level of French). These recordings, conducted as interviews, span a vast global reach—from Ivory Coast to Algeria and mainland France—making the corpus an ideal foundation for accent analysis.
 The gold-standard labels were established through a concerted labeling process. Two persons, representing different regional backgrounds (Central and Southern France) listened and annotated all recordings collaboratively,  achieving consensus after multiple passes. They had a posteriori access to the full speaker metadata (childhood and current locations at time of recording) to resolve ambiguous cases and contacted a third person in a couple of cases.
The resulting dataset features 186 annotated recordings categorized into eight distinct regional classes: Canadian French, Caribbean French, Central France, Eastern & Northern France, North Africa, Pacific and Indian Area, Southern France, Sub-Saharan Africa.