Logo Lanfrica

rjnieto/spanish-french-dialect-asr: Dialect Bias in Speech Recognition Across 10 Spanish and French Varieties

Domain:

natural language processing

Record type:

dataset
Creator:
rjn
Publisher:
Zenodo
Host:avatar
Spanish/French Dialect ASR Corpus A 20-hour, gender-balanced evaluation corpus of Spanish and French dialectal speech with manual transcriptions, built to support analysis of dialect bias in automatic speech recognition. This is the corpus accompanying "Dialect Bias in Speech Recognition Across 10 Spanish and French Varieties" (Nieto, Angelika-Nikita, Sarkis & Yang, Interspeech 2026). Contents 1,223 minutes of speech from 129 speakers across 10 country-level dialects, sourced from publicly distributed podcasts and segmented into 5-30 second single-speaker clips. Language Dialects Minutes (M/F) Speakers (M/F) Spanish Spain, Mexico, Argentina, Chile, Dominican Republic 602 (301/301) 34 (16/18) French France, Belgium, Canada, Senegal, Ivory Coast 621 (317/304) 95 (47/48) Total 10 dialects 1,223 129 Each dialect contains approximately 120 minutes of audio with at least three speakers per gender-dialect cell. Transcriptions were produced manually by authors fluent in the target languages, preserving regional orthography and dialectal markers (e.g. voseo verb forms in Spanish, regional discourse markers in French). This is deliberate: it ensures word error rate reflects model performance on authentic dialectal speech rather than normalization mismatch, and it makes grammatical divergence between varieties measurable rather than invisible. The corpus spans Europe, the Americas, and Africa, and includes both high-resource European and lower-resource African and Caribbean varieties, with finer-grained stratification (Dominican Spanish, West African French) than is typical in existing multilingual resources. Intended use This is an evaluation resource, not training data. A benchmark of this size loses its value as a benchmark if it enters model training corpora, so the terms prohibit including it in datasets, crawls, or corpora compiled for model training. Diagnostic fine-tuning for the purpose of analysis, as reported in the accompanying paper, is permitted. The terms also prohibit use for speech synthesis, voice conversion, and speaker verification. Clean single-speaker segments with speaker identifiers are usable as voice synthesis material, and this is the restriction that makes the access gate about the speakers rather than only about copyright. The prohibition does not restrict the use of speaker embeddings for dialect- or group-level analysis of the kind reported in the paper. Access The record is public; the files are restricted. Audio is available to researchers on request. To request access, use the request function on this record or email rjnieto@stanford.edu with your name, institutional affiliation, institutional email, and a brief description of your intended use. Access is granted under the corpus Terms of Access: non-commercial research and education only; no inclusion in training corpora or crawls; no synthesis, voice conversion, or speaker verification use; no redistribution; and onward sharing only to colleagues who accept the same terms. Transcriptions, metadata, model outputs, source show listing, and reproduction code are openly available at github.com. The benchmark results reported in the paper are reproducible from the released transcripts and model outputs without access to the audio; evaluating new models requires it. Rights and removal All audio derives from podcasts publicly distributed for broad consumption. The authors do not hold copyright in the source recordings and assert no open redistribution license over them. Public distribution of a podcast establishes that the material was intended to be heard. It does not establish that its creators anticipated inclusion in a speech recognition benchmark. Those are different things, and the second does not follow from the first. The shows included in this corpus are listed in the repository so that anyone whose material appears here is able to find out. If you host, appear in, or hold rights to any recording included here and would like your material removed, contact rjnieto@stanford.edu. Affected segments will be removed from subsequent versions of the corpus and current access holders will be notified. Copyright removal requests may be sent to the same address. Citation Nieto, R., Angelika-Nikita, M., Sarkis, D., & Yang, D. (2026). Dialect Bias in Speech Recognition Across 10 Spanish and French Varieties. In Proceedings of Interspeech 2026.