Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Plug-and-Play Multilingual Few-shot Spoken Words Recognition

Domaine:

natural language processing

Type de record:

papermodelsoftware
Créateur:
SaeTso
Hôte:avatar
As technology advances and digital devices become prevalent, seamless human-machine communication is increasingly gaining significance. The growing adoption of mobile, wearable, and other Internet of Things (IoT) devices has changed how we interact with these smart devices, making accurate spoken words recognition a crucial component for effective interaction. However, building robust spoken words detection system that can handle novel keywords remains challenging, especially for low-resource languages with limited training data. Here, we propose PLiX, a multilingual and plug-and-play keyword spotting system that leverages few-shot learning to harness massive real-world data and enable the recognition of unseen spoken words at test-time. Our few-shot deep models are learned with millions of one-second audio clips across 20 languages, achieving state-of-the-art performance while being highly efficient. Extensive evaluations show that PLiX can generalize to novel spoken words given as few as just one support example and performs well on unseen languages out of the box. We release models and inference code to serve as a foundation for future research and voice-enabled user interface development for emerging devices. Code: github.com

Visit

arxiv.org

Tasks

keywordsspeech processing

Tags

Audio and Speech ProcessingMachine LearningSound

Similaires

Multilingual Spoken Words CorpusmGPT: Few-Shot Learners Go MultilingualFew-shot Learning with Multilingual Language ModelsML Commons: Multilingual Spoken Words CorpusOffline Handwritten Amharic Character Recognition Using Few-shot LearningFew-Shot Multilingual Coreference Resolution Using Long-Context Large Language Models

Multilingual Spoken Words Corpus

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.

mGPT: Few-Shot Learners Go Multilingual

Recent studies report that autoregressive language models can successfully solve many NLP tasks via zero- and few-shot learning paradigms, which opens up new possibilities for using the pre-trained language models. This paper introduces two autoregressive GPT-like

Few-shot Learning with Multilingual Language Models

Large-scale generative language models such as GPT-3 are competitive few-shot learners. While these

ML Commons: Multilingual Spoken Words Corpus

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.0. The dataset contains more than 340,000 keyw

Offline Handwritten Amharic Character Recognition Using Few-shot Learning

Few-shot learning is an important, but challenging problem of machine learning aimed at learning fro

Few-Shot Multilingual Coreference Resolution Using Long-Context Large Language Models

In this work, we present our system, which ranked second in the CRAC 2025 Shared Task on Multilingua