Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Synth4Kws: Synthesized Speech for User Defined Keyword Spotting in Low Resource Environments

Domain:

natural language processing

Record type:

papermodel
Creator:
ZhuAgaBarPar
Host:avatar
One of the challenges in developing a high quality custom keyword spotting (KWS) model is the lengthy and expensive process of collecting training data covering a wide range of languages, phrases and speaking styles. We introduce Synth4Kws - a framework to leverage Text to Speech (TTS) synthesized data for custom KWS in different resource settings. With no real data, we found increasing TTS phrase diversity and utterance sampling monotonically improves model performance, as evaluated by EER and AUC metrics over 11k utterances of the speech command dataset. In low resource settings, with 50k real utterances as a baseline, we found using optimal amounts of TTS data can improve EER by 30.1% and AUC by 46.7%. Furthermore, we mix TTS data with varying amounts of real data and interpolate the real data needed to achieve various quality targets. Our experiments are based on English and single word utterances but the findings generalize to i18n languages and other keyword types. 5 pages, 5 figures, 2 tables The paper is accepted in Interspeech SynData4GenAI 2024 Workshop - syndata4genai.org

Visit

arxiv.org

Tasks

keywords

Tags

Audio and Speech ProcessingArtificial Intelligence

Similar

Low-Resource Speech Recognition and Keyword-SpottingA Deep Learning Framework for Arabic Continuous Speech Keyword Spotting in Low-Resource Settings Using Isolated-Word Keyword Spotting and Posterior Probability FunctionsTibetan-PASEM: Phonology-Aware Speech Evidence Matching for Low-Resource Tibetan Written-Query Keyword SpottingCombining tandem and hybrid systems for improved speech recognition and keyword spotting on low resource languagesFeature learning for efficient ASR-free keyword spotting in low-resource languagesLow-resource keyword spotting using contrastively trained transformer acoustic word embeddings

Low-Resource Speech Recognition and Keyword-Spotting

The IARPA Babel program ran from March 2012 to November 2016. The aim of the program was to develop

A Deep Learning Framework for Arabic Continuous Speech Keyword Spotting in Low-Resource Settings Using Isolated-Word Keyword Spotting and Posterior Probability Functions

Continuous Speech Keyword Spotting (CSKWS) presents a challenging paradigm shift from isolated-word

Tibetan-PASEM: Phonology-Aware Speech Evidence Matching for Low-Resource Tibetan Written-Query Keyword Spotting

Low-resource written-query keyword spotting detects a text-specified target in speech without spoken

Combining tandem and hybrid systems for improved speech recognition and keyword spotting on low resource languages

Copyright © 2014 ISCA. In recent years there has been significant interest in Automatic Speech Recog

Feature learning for efficient ASR-free keyword spotting in low-resource languages

We consider feature learning for efficient keyword spotting that can be applied in severely under-re

Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings

We introduce a new approach, the ContrastiveTransformer, that produces acoustic word embeddings (AWE