Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Comparing decoding strategies for subword-based keyword spotting in low-resourced languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
HarLe,MesLam
Éditeur:
LabVocISCAISCA
Éditeur:
CCSD
Hôte:avatar
International audience For languages with limited training resources, out-of-vocabulary (OOV) words are a significant problem, both fortranscription and keyword spotting. This paper investigates theuse of subword lexical units for keyword spotting. Three strate-gies for using the sub-word units are explored: 1) convertingword-based lattices to subword lattices after decoding, 2) per-forming a separate decoding for each subword type, and 3) asingle decoding using all possible subword units. In these ex-periments, the best performance is achieved by carrying out aseparate decoding for each subword type. Further gains are at-tained through system combination. We also find that ignor-ing word boundaries improves the detection of OOV keywordswithout significantly impacting in-vocabulary keyword detec-tion. Results are presented on four languages from the IARPABabel Program (Haitian Creole, Assamese, Bengali, and Zulu).

Visit

hal.science

Tasks

keywordsspeech processing

Tags

keyword searchspoken term detectionOOVsub-word lexical unitslow resource LVCSR[INFO]Computer Science [cs][INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]