Logo Lanfrica

kamperh/bucktsong_segmentalist

Domain:

natural language processing
Creator:
kam
Host:
Unsupervised segmentation and clustering of Buckeye English and NCHLT Xitsonga corpora. Recipe: Segmentation and Clustering of Buckeye English and NCHLT Xitsonga ========================================================================= Contributors ------------ - Herman Kamper - Aren Jansen - Sharon Goldwater Overview -------- This is a recipe for unsupervised segmentation and clustering of subsets of the Buckeye English and NCHLT Xitsonga corpora. Details of the approach is given in Kamper et al., 2016: - H. Kamper, A. Jansen, and S. J. Goldwater, "A segmental framework for fully-unsupervised large-vocabulary speech recognition," *arXiv preprint arXiv:1606.06950*, 2016. Please cite this paper if you use this code. The recipe below makes use of the separate segmentalist package which performs the actual unsupervised segmentation and clustering and was developed together with this recipe. Disclaimer ---------- The code provided here is not pretty. But I believe that research should be reproducible, and I hope that this repository is sufficient to make this possible for the paper mentioned above. I provide no guarantees with the code, but please let me know if you have any problems, find bugs or have general comments. Datasets -------- Portions of the Buckeye English and NCHLT Xitsonga corpora are used. The whole Buckeye corpus will be required to execute the steps here, and the portion of the NCHLT data. These can be downloaded from: - Buckeye corpus: buckeyecorpus.osu.edu - NCHLT Xitsonga portion: www.zerospeech.com. This requires registration for the challenge. From the complete Buckeye corpus we split off several subsets. The most important are the sets labelled as `devpart1` and `zs` in the code here. These sets respectively correspond to `English1` and `English2` in Kamper et al., 2016, so see the paper for more details. More details of which speakers are found in which set is also given at the end of features/readme.md. We use the entire Xitsonga dataset provided as part of the Zero Speech Challenge 2015 (this was already a subset o …