Logo Lanfrica

SpeechColab/GigaSpeech2

Domaine:

natural language processing

Type de record:

datasetmodel
Créateur:
Spe
Hôte:
An evolving, large-scale and multi-domain ASR corpus for low-resource languages with automated crawling, transcription and refinement # GigaSpeech 2 This is the official repository of the GigaSpeech 2 dataset. For details of how we created the dataset, please refer to our ACL 2025 Main paper. GigaSpeech 2 version: 2.0 (2024/06/19) ## Download * The dataset is available at HuggingFace and ModelScope. * The pre-trained models are available at Thai 10000h and Vietnamese 70000h. ## Leaderboard | **Contributor**| **Toolkit** | **Train Recipe** | **Train Data** |**Test CER/WER** | |:---------------|:------------------|:------------------|:------------------|:------------------| ||||| | Baseline | Icefall | Zipformer/Stateless pruned RNN-T | GigaSpeech 2.0 th | 12.46 | | Baseline | Icefall | Zipformer/Stateless pruned RNN-T | GigaSpeech 2.0 id | 14.92 | | Baseline | Icefall | Zipformer/Stateless pruned RNN-T | GigaSpeech 2.0 vi | 12.83 | | Baseline | ESPNet | Conformer/Transformer CTC/AED | GigaSpeech 2.0 th | 13.70 | | Baseline | ESPNet | Conformer/Transformer CTC/AED | GigaSpeech 2.0 id | 15.50 | | Baseline | ESPNet | Conformer/Transformer CTC/AED | GigaSpeech 2.0 vi | 14.60 | ## Dataset ### Audio Source * Language: Thai, Indonesian, Vietnamese * GigaSpeech 2 raw: 30,000 hours of automatically transcribed speech across Thai, Indonesian, and Vietnamese. * GigaSpeech 2 refined: 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. * GigaSpeech 2 DEV & TEST: 10 hours for DEV and 10 hours for TEST per language, **transcribed by professional human annotators**, challenging and realistic. ### Training Subsets | | Thai (hours) | Indonesian (hours) | Vietnamese (hours) | |:--------------------:|:------------:|:------------------:|:------------------:| | GigaSpeech 2 raw | 12901.8 | 8112.9 | 7324.0 | | GigaSpeech 2 refined | 10262.0 | 5714.0 | 6039.0 | GigaSpeech 2 raw contains all the data from GigaSpeech 2 refined. ### Evaluation Subsets | …