Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings

Domaine:

natural language processing

Type de record:

paper
Créateur:
HerJacNiesler, Thomas
Hôte:avatar
We introduce a new approach, the ContrastiveTransformer, that produces acoustic word embeddings (AWEs) for the purpose of very low-resource keyword spotting. The ContrastiveTransformer, an encoder-only model, directly optimises the embedding space using normalised temperature-scaled cross entropy (NT-Xent) loss. We use this model to perform keyword spotting for radio broadcasts in Luganda and Bambara, the latter a severely under-resourced language. We compare our model to various existing AWE approaches, including those constructed from large pre-trained self-supervised models, a recurrent encoder which previously used the NT-Xent loss, and a DTW baseline. We demonstrate that the proposed contrastive transformer approach offers performance improvements over all considered existing approaches to very low-resource keyword spotting in both languages. 5 pages, 2 figures

Visit

arxiv.org

Tasks

keywordsspeech processing

Languages

BamanankanGandaLame

Tags

Audio and Speech Processing

Similaires

A Deep Learning Framework for Arabic Continuous Speech Keyword Spotting in Low-Resource Settings Using Isolated-Word Keyword Spotting and Posterior Probability FunctionsMultilingual Jointly Trained Acoustic and Written Word EmbeddingsLow-Resource Speech Recognition and Keyword-SpottingAnalyzing Acoustic Word Embeddings from Pre-trained Self-supervised ModelsAnalyzing Acoustic Word Embeddings from Pre-trained Self-supervised Speech ModelsEvaluating Performance of Pre-trained Word Embeddings on Assamese, a Low-resource Language

A Deep Learning Framework for Arabic Continuous Speech Keyword Spotting in Low-Resource Settings Using Isolated-Word Keyword Spotting and Posterior Probability Functions

Continuous Speech Keyword Spotting (CSKWS) presents a challenging paradigm shift from isolated-word

Multilingual Jointly Trained Acoustic and Written Word Embeddings

Acoustic word embeddings (AWEs) are vector representations of spoken word segments. AWEs can be lear

Low-Resource Speech Recognition and Keyword-Spotting

The IARPA Babel program ran from March 2012 to November 2016. The aim of the program was to develop

Analyzing Acoustic Word Embeddings from Pre-trained Self-supervised Models

IEEE ICASSP 2023 Conference, Hybrid Event, 4-10 June 2023, Rhodes Island, Greece Given the strong re

Analyzing Acoustic Word Embeddings from Pre-trained Self-supervised Speech Models

Given the strong results of self-supervised models on various tasks, there have been surprisingly fe

Evaluating Performance of Pre-trained Word Embeddings on Assamese, a Low-resource Language