
This dataset is a large-scale multilingual speech corpus curated for speech-to-speech translation, speech-to-text, and multilingual speech processing research. The data is organized by language, speaker (user_id), and dataset split (train, dev), and includes rich acoustic and metadata annotations.
The dataset is published and maintained by Marie Maltais at mcgill-NLP.