Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Unsupervised Acoustic Unit Discovery by Leveraging a Language-Independent Subword Discriminative Feature Representation

Domaine:

natural language processing

Type de record:

paper
Créateur:
FenŻelMorSch
Hôte:avatar
This paper tackles automatically discovering phone-like acoustic units (AUD) from unlabeled speech data. Past studies usually proposed single-step approaches. We propose a two-stage approach: the first stage learns a subword-discriminative feature representation and the second stage applies clustering to the learned representation and obtains phone-like clusters as the discovered acoustic units. In the first stage, a recently proposed method in the task of unsupervised subword modeling is improved by replacing a monolingual out-of-domain (OOD) ASR system with a multilingual one to create a subword-discriminative representation that is more language-independent. In the second stage, segment-level k-means is adopted, and two methods to represent the variable-length speech segments as fixed-dimension feature vectors are compared. Experiments on a very low-resource Mboshi language corpus show that our approach outperforms state-of-the-art AUD in both normalized mutual information (NMI) and F-score. The multilingual ASR improved upon the monolingual ASR in providing OOD phone labels and in estimating the phone boundaries. A comparison of our systems with and without knowing the ground-truth phone boundaries showed a 16% NMI performance gap, suggesting that the current approach can significantly benefit from improved phone boundary estimation. Accepted for publication in INTERSPEECH 2021

Visit

arxiv.org

Tasks

speech processing

Languages

Mbosi

Tags

Audio and Speech ProcessingComputation and LanguageSound

Similaires

A Hierarchical Subspace Model for Language-Attuned Acoustic Unit DiscoveryLanguage independent and unsupervised acoustic models for speech recognition and keyword spottingVoice Conversion Based Speaker Normalization for Acoustic Unit DiscoveryNon-Parametric Bayesian Subspace Models for Acoustic Unit DiscoveryUnsupervised word discovery for computational language documentationBayesian Models for Unit Discovery on a Very Low Resource Language

A Hierarchical Subspace Model for Language-Attuned Acoustic Unit Discovery

In this work, we propose a hierarchical subspace model for acoustic unit discovery. In this approach

Language independent and unsupervised acoustic models for speech recognition and keyword spotting

Copyright © 2014 ISCA. Developing high-performance speech processing systems for low-resource langua

Voice Conversion Based Speaker Normalization for Acoustic Unit Discovery

Discovering speaker independent acoustic units purely from spoken input is known to be a hard proble

Non-Parametric Bayesian Subspace Models for Acoustic Unit Discovery

This work investigates subspace non-parametric models for the task of learning a set of acoustic uni

Unsupervised word discovery for computational language documentation

Découverte non-supervisée de mots pour outiller la linguistique de terrain La dive

Bayesian Models for Unit Discovery on a Very Low Resource Language

Developing speech technologies for low-resource languages has become a very active research field ov