Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Comparing decoding strategies for subword-based keyword spotting in low-resourced languages

Domain:

natural language processing

Record type:

paper
Creator:
HarLe,MesLam
Editor:
LabVocISCAISCA
Publisher:
CCSD
Host:avatar
International audience For languages with limited training resources, out-of-vocabulary (OOV) words are a significant problem, both fortranscription and keyword spotting. This paper investigates theuse of subword lexical units for keyword spotting. Three strate-gies for using the sub-word units are explored: 1) convertingword-based lattices to subword lattices after decoding, 2) per-forming a separate decoding for each subword type, and 3) asingle decoding using all possible subword units. In these ex-periments, the best performance is achieved by carrying out aseparate decoding for each subword type. Further gains are at-tained through system combination. We also find that ignor-ing word boundaries improves the detection of OOV keywordswithout significantly impacting in-vocabulary keyword detec-tion. Results are presented on four languages from the IARPABabel Program (Haitian Creole, Assamese, Bengali, and Zulu).

Visit

hal.science

Tasks

keywordsspeech processing

Tags

keyword searchspoken term detectionOOVsub-word lexical unitslow resource LVCSR[INFO]Computer Science [cs][INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]

Similar

Joint decoding of tandem and hybrid systems for improved keyword spotting on low resource languagesLanguage Model Data Augmentation for Keyword Spotting in Low-Resourced Training ConditionsFeature learning for efficient ASR-free keyword spotting in low-resource languagesKeyword Spotting with African LanguagesLow-Resource Speech Recognition and Keyword-SpottingCombining tandem and hybrid systems for improved speech recognition and keyword spotting on low resource languages

Joint decoding of tandem and hybrid systems for improved keyword spotting on low resource languages

Copyright © 2015 ISCA. Keyword spotting (KWS) for low-resource languages has drawn increasing attent

Language Model Data Augmentation for Keyword Spotting in Low-Resourced Training Conditions

International audience

Feature learning for efficient ASR-free keyword spotting in low-resource languages

We consider feature learning for efficient keyword spotting that can be applied in severely under-re

Keyword Spotting with African Languages

Keyword spotting refers to the task of learning to detect spoken keywords. It

Low-Resource Speech Recognition and Keyword-Spotting

The IARPA Babel program ran from March 2012 to November 2016. The aim of the program was to develop

Combining tandem and hybrid systems for improved speech recognition and keyword spotting on low resource languages

Copyright © 2014 ISCA. In recent years there has been significant interest in Automatic Speech Recog