Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically

Domaine:

natural language processing

Type de record:

paper
Créateur:
ShiDe Hu,Vie
Hôte:avatar
Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on spoken translation retrieval using representational similarity. However, prior work does not control for phonetic overlap between equivalent utterances, which may artificially support retrieval. We conduct pronunciation-controlled experiments to test whether cross-lingual alignment arises from semantic rather than phonetic similarity. Results show that spoken translation retrieval remains strongly above chance without phonetic cues in the final layers of encoders trained with a speech translation objective, most clearly for models additionally trained on translation. We further test early-exiting the encoder to induce representations we hypothesize to be less tied to language-specific semantics. These experiments indeed reveal performance gains in automatic speech recognition on low-resource languages unseen during training. Submitted to Interspeech 2026

Visit

arxiv.org

Tasks

automatic speech recognitionmachine translationspeech processingspeech translation

Tags

Computation and Language

Similaires

POWSM: A Phonetic Open Whisper-Style Speech Foundation ModelSemantically Corrected Amharic Automatic Speech RecognitionPerformance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian DialectOperationalizing Style: Quantifying the Use of Style Shift in the Speech of African American AdolescentsJoe254h/whisper-wolof-speechTusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition Experiments

POWSM: A Phonetic Open Whisper-Style Speech Foundation Model

Recent advances in spoken language processing have led to substantial progress in phonetic tasks suc

Semantically Corrected Amharic Automatic Speech Recognition

Automatic Speech Recognition (ASR) can play a crucial role in enhancing the accessibility of spoken

Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect

Speech encoders pretrained through self-supervised learning (SSL) have demonstrated remarkable perfo

Operationalizing Style: Quantifying the Use of Style Shift in the Speech of African American Adolescents

The vast majority of research to date on African American Vernacular English style shift has taken t

Joe254h/whisper-wolof-speech

Flask web app for Wolof speech transcription using a fine-tuned Whisper model. --- title: Wolof Whi

Tusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition Experiments

There is growing interest in ASR systems that can recognize phones in a language-independent fashion