Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Additional Shared Decoder on Siamese Multi-view Encoders for Learning Acoustic Word Embeddings

Domain:

natural language processing

Record type:

paper
Creator:
JunLimGooJun
Host:avatar
Acoustic word embeddings --- fixed-dimensional vector representations of arbitrary-length words --- have attracted increasing interest in query-by-example spoken term detection. Recently, on the fact that the orthography of text labels partly reflects the phonetic similarity between the words' pronunciation, a multi-view approach has been introduced that jointly learns acoustic and text embeddings. It showed that it is possible to learn discriminative embeddings by designing the objective which takes text labels as well as word segments. In this paper, we propose a network architecture that expands the multi-view approach by combining the Siamese multi-view encoders with a shared decoder network to maximize the effect of the relationship between acoustic and text embeddings in embedding space. Discriminatively trained with multi-view triplet loss and decoding loss, our proposed approach achieves better performance on acoustic word discrimination task with the WSJ dataset, resulting in 11.1% relative improvement in average precision. We also present experimental results on cross-view word discrimination and word level speech recognition tasks. Accepted at 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2019)

Visit

arxiv.org

Tasks

automatic speech recognitionembeddingsspeech processing

Tags

Audio and Speech ProcessingInformation RetrievalMachine LearningSound

Similar

Multilingual acoustic word embeddings for zero-resource languagesMultilingual Jointly Trained Acoustic and Written Word EmbeddingsAnalyzing Acoustic Word Embeddings from Pre-trained Self-supervised ModelsSHAPED: Shared-Private Encoder-Decoder for Text Style AdaptationAnalyzing Acoustic Word Embeddings from Pre-trained Self-supervised Speech ModelsLow-resource keyword spotting using contrastively trained transformer acoustic word embeddings

Multilingual acoustic word embeddings for zero-resource languages

This research addresses the challenge of developing speech applications for zero-resource languages

Multilingual Jointly Trained Acoustic and Written Word Embeddings

Acoustic word embeddings (AWEs) are vector representations of spoken word segments. AWEs can be lear

Analyzing Acoustic Word Embeddings from Pre-trained Self-supervised Models

IEEE ICASSP 2023 Conference, Hybrid Event, 4-10 June 2023, Rhodes Island, Greece Given the strong re

SHAPED: Shared-Private Encoder-Decoder for Text Style Adaptation

Supervised training of abstractive language generation models results in learning conditional probab

Analyzing Acoustic Word Embeddings from Pre-trained Self-supervised Speech Models

Given the strong results of self-supervised models on various tasks, there have been surprisingly fe

Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings

We introduce a new approach, the ContrastiveTransformer, that produces acoustic word embeddings (AWE