Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Unsupervised feature learning for speech using correspondence and Siamese networks

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
LasEngKam
Hôte:avatar
In zero-resource settings where transcribed speech audio is unavailable, unsupervised feature learning is essential for downstream speech processing tasks. Here we compare two recent methods for frame-level acoustic feature learning. For both methods, unsupervised term discovery is used to find pairs of word examples of the same unknown type. Dynamic programming is then used to align the feature frames between each word pair, serving as weak top-down supervision for the two models. For the correspondence autoencoder (CAE), matching frames are presented as input-output pairs. The Triamese network uses a contrastive loss to reduce the distance between frames of the same predicted word type while increasing the distance between negative examples. For the first time, these feature extractors are compared on the same discrimination tasks using the same weak supervision pairs. We find that, on the two datasets considered here, the CAE outperforms the Triamese network. However, we show that a new hybrid correspondence-Triamese approach (CTriamese), consistently outperforms both the CAE and Triamese models in terms of average precision and ABX error rates on both English and Xitsonga evaluation data. 5 pages, 3 figures, 2 tables; accepted to the IEEE Signal Processing Letters, (c) 2020 IEEE

Visit

arxiv.org

Tasks

speech processing

Languages

Tsonga

Tags

Computation and LanguageAudio and Speech Processing

Similaires

Unsupervised learning for expressive speech synthesisStock feature dimensionality reduction for closing price prediction using unsupervised machine learning technique (case study of Nigeria Stock Exchange)Learning English and Arabic Question Similarity with Siamese Neural Networks in Community Question Answering servicesSecuring smart agriculture networks using bio-inspired feature selection and transfer learning for effective image-based intrusion detectionFeature exploration for almost zero-resource ASR-free keyword spotting using a multilingual bottleneck extractor and correspondence autoencodersUnsupervised Domain Adaptation with Feature Embeddings

Unsupervised learning for expressive speech synthesis

Nowadays, especially with the upswing of neural networks, speech synthesis is almost totally data dr

Stock feature dimensionality reduction for closing price prediction using unsupervised machine learning technique (case study of Nigeria Stock Exchange)

Stock data offers invaluable insights into the world of finance. It encourages investment and saving

Learning English and Arabic Question Similarity with Siamese Neural Networks in Community Question Answering services

International audience In this paper, we tackle the task of similar question retrieva

Securing smart agriculture networks using bio-inspired feature selection and transfer learning for effective image-based intrusion detection

Feature exploration for almost zero-resource ASR-free keyword spotting using a multilingual bottleneck extractor and correspondence autoencoders

We compare features for dynamic time warping (DTW) when used to bootstrap keyword spotting (KWS) in an almost zero-resource setting. Such quickly-deployable systems aim to support United Nations (UN) humanitarian relief efforts in parts of Africa with severely unde

Unsupervised Domain Adaptation with Feature Embeddings

Representation learning is the dominant technique for unsupervised domain adaptation, but existing a