This repo contains code used to implement models used in our INTERSPEECH 2021 paper:
S. Feng, P. \.Zelasko, L. Moro{-}Vel{\'{a}}zquez and O. Scharenborg, "Unsupervised Acoustic Unit Discovery by Leveraging a Language-independent Subword Discriminative Feature Representation", in Proc. INTERSPEECH 2021. Link:
isca-speech.org
Update on 21 Sep
- Support phoneme NMI evaluation on Mboshi. See s5/scripts/run/disco_multi_tdnnf_1g_own_gmm.sh.
Update on 16 Sep
- Support phone2phoneme mapping for Mboshi. See s5/references/ and s5/scripts/run/phone2phoneme_conversion.py
Task:
Unsupervised acoustic unit discovery
Databases:
Evaluation: Mboshi, freely available for academic use.
Training: Globalphone 5 languages (Czech, French, Spanish, Mandarin and Thai), 8 Babel languages (Cantonese, Bengali, Vietnamese, Lao, Zulu, Amharic, Javanese and Georgian). Please note that GlobalPhone and IARPA BABEL Language resources are *NOT* freely available. Also note that the training of phone-level multilingual ASR systems with 5 or 13 GP & Babel languages is not included here, please refer to
github.com and paper: Feng et al. ICASSP 2021 "HOW PHONOTACTICS AFFECT MULTILINGUAL AND ZERO-SHOT ASR PERFORMANCE" for details of building phone-level (IPA) multilingual ASR systems.
Evaluation metrics:
(1) Normalized mutual information (NMI): a measure of clustering quality
(2) F-score: a measure of segmentation quality
Please refer to
github.com ([1]) for evaluation software.
See our Interspeech 2021 paper for details of the approach.
References
[1] B. Yusuf, L. Ondel, L. Burget, J. Cernock ́y, and M. Saraclar, “Ahierarchical subspace model for language-attuned acoustic unitdiscovery,” CoRR, vol. abs/2011.03115, 2020.