Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Speechless: Speech Instruction Training Without Speech for Low Resource Languages

Domaine:

natural language processing

Type de record:

papermodelsoftware
Créateur:
DaoVu,Ha,Anh
Hôte:avatar
The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of speech recognition data, there is a notable scarcity of speech instruction data, which is essential for fine-tuning models to understand and execute spoken commands. Generating high-quality synthetic speech requires a good text-to-speech (TTS) model, which may not be available to low resource languages. Our novel approach addresses this challenge by halting synthesis at the semantic representation level, bypassing the need for TTS. We achieve this by aligning synthetic semantic representations with the pre-trained Whisper encoder, enabling an LLM to be fine-tuned on text instructions while maintaining the ability to understand spoken instructions during inference. This simplified training process is a promising approach to building voice assistant for low-resource languages. This paper was accepted by INTERSPEECH 2025

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingComputation and LanguageSound

Similaires

Robust speech recognition for low-resource languagesEnhancing Automatic Speech Recognition for Child Speech in Low-Resource LanguagesSANTLR: Speech Annotation Toolkit for Low Resource LanguagesAdversarial Text-to-Speech for low-resource languagesDEFI-COLaF/Speech-Recognition-for-Low-Resource-LanguagesDataScience-ArtificialIntelligence/Hate-speech-detection-for-Low-resource-languages

Robust speech recognition for low-resource languages

Process of human-machine interaction is an integral part of everyday human life in a modern world. T

Enhancing Automatic Speech Recognition for Child Speech in Low-Resource Languages

Automatic speech recognition (ASR) for children is demanding because their speech differs c

SANTLR: Speech Annotation Toolkit for Low Resource Languages

While low resource speech recognition has attracted a lot of attention from the speech community, th

Adversarial Text-to-Speech for low-resource languages

Improving the adversarial TTS models for low-resource languages by utilizing the high-frequency similarities between the different languages.

DEFI-COLaF/Speech-Recognition-for-Low-Resource-Languages

# Speech-Recognition-for-Low-Resource-Languages This repository contains code to fine-tune Whisper

DataScience-ArtificialIntelligence/Hate-speech-detection-for-Low-resource-languages

# Hate-speech-detection-for-Low-resource-languages ## Overview This project aims to develop machine