This repository is being prepared as a row-normalized multilingual
ASR dataset built from Mozilla Data Collective Common Voice
Scripted Speech 25.0. The normalized rows are the main deliverable:
audio, sentence, locale, language, split,
source_dataset_id, source_archive, and upstream Common Voice
metadata where present.
During staging, the original MDC .tar.gz archives are preserved