Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning

Domain:

natural language processing

Record type:

paper
Creator:
LauChen, QianFanXu,
Host:avatar
Our quality audit for three widely used public multilingual speech datasets - Mozilla Common Voice 17.0, FLEURS, and Vox Populi - shows that in some languages, these datasets suffer from significant quality issues, which may obfuscate downstream evaluation results while creating an illusion of success. We divide these quality issues into two categories: micro-level and macro-level. We find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages. We provide a case analysis of Taiwanese Southern Min (nan_tw) that highlights the need for proactive language planning (e.g. orthography prescriptions, dialect boundary definition) and enhanced data quality control in the dataset creation process. We conclude by proposing guidelines and recommendations to mitigate these issues in future dataset development, emphasizing the importance of sociolinguistic awareness and language planning principles. Furthermore, we encourage research into how this creation process itself can be leveraged as a tool for community-led language planning and revitalization. Accepted by ACL 2025 Main Conference

Visit

arxiv.org

Tasks

speech processing

Tags

Computation and LanguageArtificial Intelligence

Similar

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African LanguagesDeveloping Nigeria Multilingual Languages Speech Datasets for Antenatal OrientationIssues in the development of cross-cultural assessments of speech and language for childrenWaxal-Multilingual/speech-dataIntersectional Bias in Hate Speech and Abusive Language DatasetsLANGUAGE-IN-EDUCATION PLANNING IN TANZANIA: A SOCIOLINGUISTIC ANALYSIS

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

Hate speech and abusive language are global phenomena that need socio-cultural background knowledge

Developing Nigeria Multilingual Languages Speech Datasets for Antenatal Orientation

Issues in the development of cross-cultural assessments of speech and language for children

Background: There is an increasing demand for the assessment of speech and langu

Waxal-Multilingual/speech-data

This repository contains multi-modal speech data for African languages that can be used to train ASR

Intersectional Bias in Hate Speech and Abusive Language Datasets

Algorithms are widely applied to detect hate speech and abusive language in social media. We investi

LANGUAGE-IN-EDUCATION PLANNING IN TANZANIA: A SOCIOLINGUISTIC ANALYSIS