Unsupervised segmentation and clustering of Buckeye English and NCHLT Xitsonga corpora.
Recipe: Segmentation and Clustering of Buckeye English and NCHLT Xitsonga
=========================================================================
Contributors
------------
- Herman Kamper
- Aren Jansen
- Sharon Goldwater
Overview
--------
This is a recipe for unsupervised segmentation and clustering of subsets of the
Buckeye English and NCHLT Xitsonga corpora. Details of the approach is given in
Kamper et al., 2016:
- H. Kamper, A. Jansen, and S. J. Goldwater, "A segmental framework for
fully-unsupervised large-vocabulary speech recognition," *arXiv preprint
arXiv:1606.06950*, 2016.
Please cite this paper if you use this code.
The recipe below makes use of the separate
segmentalist package which performs
the actual unsupervised segmentation and clustering and was developed together
with this recipe.
Disclaimer
----------
The code provided here is not pretty. But I believe that research should be
reproducible, and I hope that this repository is sufficient to make this
possible for the paper mentioned above. I provide no guarantees with the code,
but please let me know if you have any problems, find bugs or have general
comments.
Datasets
--------
Portions of the Buckeye English and NCHLT Xitsonga corpora are used. The whole
Buckeye corpus will be required to execute the steps here, and the portion of
the NCHLT data. These can be downloaded from:
- Buckeye corpus:
buckeyecorpus.osu.edu
- NCHLT Xitsonga portion: www.zerospeech.com. This requires registration
for the challenge.
From the complete Buckeye corpus we split off several subsets. The most
important are the sets labelled as `devpart1` and `zs` in the code here. These
sets respectively correspond to `English1` and `English2` in Kamper et al.,
2016, so see the paper for more details. More
details of which speakers are found in which set is also given at the end of
features/readme.md. We use the entire Xitsonga dataset
provided as part of the Zero Speech Challenge 2015 (this was already a subset
o …