An evolving, large-scale and multi-domain ASR corpus for low-resource languages with automated crawling, transcription and refinement
# GigaSpeech 2
This is the official repository of the GigaSpeech 2 dataset. For details of how we created the dataset, please refer to our ACL 2025 Main paper.
GigaSpeech 2 version: 2.0 (2024/06/19)
## Download
* The dataset is available at HuggingFace and ModelScope.
* The pre-trained models are available at Thai 10000h and Vietnamese 70000h.
## Leaderboard
| **Contributor**| **Toolkit** | **Train Recipe** | **Train Data** |**Test CER/WER** |
|:---------------|:------------------|:------------------|:------------------|:------------------|
|||||
| Baseline | Icefall | Zipformer/Stateless pruned RNN-T | GigaSpeech 2.0 th | 12.46 |
| Baseline | Icefall | Zipformer/Stateless pruned RNN-T | GigaSpeech 2.0 id | 14.92 |
| Baseline | Icefall | Zipformer/Stateless pruned RNN-T | GigaSpeech 2.0 vi | 12.83 |
| Baseline | ESPNet | Conformer/Transformer CTC/AED | GigaSpeech 2.0 th | 13.70 |
| Baseline | ESPNet | Conformer/Transformer CTC/AED | GigaSpeech 2.0 id | 15.50 |
| Baseline | ESPNet | Conformer/Transformer CTC/AED | GigaSpeech 2.0 vi | 14.60 |
## Dataset
### Audio Source
* Language: Thai, Indonesian, Vietnamese
* GigaSpeech 2 raw: 30,000 hours of automatically transcribed speech across Thai, Indonesian, and Vietnamese.
* GigaSpeech 2 refined: 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese.
* GigaSpeech 2 DEV & TEST: 10 hours for DEV and 10 hours for TEST per language, **transcribed by professional human annotators**, challenging and realistic.
### Training Subsets
| | Thai (hours) | Indonesian (hours) | Vietnamese (hours) |
|:--------------------:|:------------:|:------------------:|:------------------:|
| GigaSpeech 2 raw | 12901.8 | 8112.9 | 7324.0 |
| GigaSpeech 2 refined | 10262.0 | 5714.0 | 6039.0 |
GigaSpeech 2 raw contains all the data from GigaSpeech 2 refined.
### Evaluation Subsets
| …