Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

That Sounds Familiar: an Analysis of Phonetic Representations Transfer Across Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
ŻelMorHasSch
Hôte:avatar
Only a handful of the world's languages are abundant with the resources that enable practical applications of speech processing technologies. One of the methods to overcome this problem is to use the resources existing in other languages to train a multilingual automatic speech recognition (ASR) model, which, intuitively, should learn some universal phonetic representations. In this work, we focus on gaining a deeper understanding of how general these representations might be, and how individual phones are getting improved in a multilingual setting. To that end, we select a phonetically diverse set of languages, and perform a series of monolingual, multilingual and crosslingual (zero-shot) experiments. The ASR is trained to recognize the International Phonetic Alphabet (IPA) token sequences. We observe significant improvements across all languages in the multilingual setting, and stark degradation in the crosslingual setting, where the model, among other errors, considers Javanese as a tone language. Notably, as little as 10 hours of the target language training data tremendously reduces ASR error rates. Our analysis uncovered that even the phones that are unique to a single language can benefit greatly from adding training data from other languages - an encouraging result for the low-resource speech community. Submitted to Interspeech 2020. For some reason, the ArXiv Latex engine rendered it in more than 4 pages

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingComputation and LanguageSound

Similaires

MURAL: Multimodal, Multitask Representations Across LanguagesNLP-based Representations of Sentiment Analysis for African LanguagesAnalyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal RepresentationsWitty-Kitty/Phoneme-Recognition-through-Fine-Tuning-of-Phonetic-RepresentationsSawtone: A universal framework for phonetic similarity and alignment across languages and scriptsA Phonetic Study of West African Languages - An auditory-instrumental survey

MURAL: Multimodal, Multitask Representations Across Languages

Both image-caption pairs and translation pairs provide the means to learn deep representations of an

NLP-based Representations of Sentiment Analysis for African Languages

This comparison gives representations of sentiment analysis for African languages, the various metho

Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations

While understanding the knowledge boundaries of LLMs is crucial to prevent hallucination, research o

Witty-Kitty/Phoneme-Recognition-through-Fine-Tuning-of-Phonetic-Representations

This repository makes available code for running expetiments to finetune Allosaurus on Bukusu and Sa

Sawtone: A universal framework for phonetic similarity and alignment across languages and scripts

Processing text across different scripts presents significant hurdles in natural language processing

A Phonetic Study of West African Languages - An auditory-instrumental survey