Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A comparison of self-supervised speech representations as input features for unsupervised acoustic word embeddings

Domain:

natural language processing

Record type:

paper
Creator:
vanKam
Host:avatar
Many speech processing tasks involve measuring the acoustic similarity between speech segments. Acoustic word embeddings (AWE) allow for efficient comparisons by mapping speech segments of arbitrary duration to fixed-dimensional vectors. For zero-resource speech processing, where unlabelled speech is the only available resource, some of the best AWE approaches rely on weak top-down constraints in the form of automatically discovered word-like segments. Rather than learning embeddings at the segment level, another line of zero-resource research has looked at representation learning at the short-time frame level. Recent approaches include self-supervised predictive coding and correspondence autoencoder (CAE) models. In this paper we consider whether these frame-level features are beneficial when used as inputs for training to an unsupervised AWE model. We compare frame-level features from contrastive predictive coding (CPC), autoregressive predictive coding and a CAE to conventional MFCCs. These are used as inputs to a recurrent CAE-based AWE model. In a word discrimination task on English and Xitsonga data, all three representation learning approaches outperform MFCCs, with CPC consistently showing the biggest improvement. In cross-lingual experiments we find that CPC features trained on English can also be transferred to Xitsonga. Accepted to SLT 2021

Visit

arxiv.org

Tasks

embeddingsspeech processing

Languages

Tsonga

Tags

Computation and LanguageAudio and Speech Processing

Similar

Analyzing Acoustic Word Embeddings from Pre-trained Self-supervised Speech ModelsAnalyzing Acoustic Word Embeddings from Pre-trained Self-supervised ModelsPhoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment EvaluationSelf-Supervised Acoustic Word Embedding Learning via Correspondence Transformer EncoderDo Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech

Analyzing Acoustic Word Embeddings from Pre-trained Self-supervised Speech Models

Given the strong results of self-supervised models on various tasks, there have been surprisingly fe

Analyzing Acoustic Word Embeddings from Pre-trained Self-supervised Models

IEEE ICASSP 2023 Conference, Hybrid Event, 4-10 June 2023, Rhodes Island, Greece Given the strong re

Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation

Forced alignment evaluation typically requires manually annotated timestamps, limiting large-scale a

Self-Supervised Acoustic Word Embedding Learning via Correspondence Transformer Encoder

Acoustic word embeddings (AWEs) aims to map a variable-length speech segment into a fixed-dimensiona

Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?

Discrete representations of speech, obtained from Self-Supervised Learning (SSL) foundation models,

Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech

Annotating a multilingual code-switched corpus is a painstaking process requiring specialist linguis