This dataset is an unofficial version of the Mozilla Common Voice Corpus 16. It was downloaded and c
The Common Voice dataset consists of a unique MP3 and corresponding text file. Many of the 30328 rec
This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and c
The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identi
Common Voice is Mozilla's initiative to help teach machines how real people speak. The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.
Voice data collection and distribution Interface for Basaa language using Mozilla's Common Voice infrastructure.