Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

An Empirical Recipe for Universal Phone Recognition

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
BhaLi,ChoYeo
Hôte:avatar
Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive. Highly performant English-focused models do not generalize across languages, while multilingual models underutilize pretrained representations. It also remains unclear how data scale, architecture, and training objective contribute to multilingual PR. We present PhoneticXEUS -- trained on large-scale multilingual data and achieving state-of-the-art performance on both multilingual (17.7% PFER) and accented English speech (10.6% PFER). Through controlled ablations with evaluations across 100+ languages under a unified scheme, we empirically establish our training recipe and quantify the impact of SSL representations, data scale, and loss objectives. In addition, we analyze error patterns across language families, accented speech, and articulatory features. All data and code are released openly. Submitted to Interspeech 2026. Code: github.com

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageMachine LearningSoundAudio and Speech Processing

Similaires

Tusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition ExperimentsUniversal Phone Recognition with a Multilingual Allophone SystemPhone inventory optimization for multilingual automatic speech recognitionTowards Universal Khmer Text RecognitionBankNote-Net: Open dataset for assistive universal currency recognitionUsing articulatory feature detectors in progressive networks for multilingual low-resource phone recognition

Tusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition Experiments

There is growing interest in ASR systems that can recognize phones in a language-independent fashion

Universal Phone Recognition with a Multilingual Allophone System

Multilingual models can improve language processing, particularly for low resource situations, by sh

Phone inventory optimization for multilingual automatic speech recognition

This paper describes a phone inventory optimization procedure for application in multilingual automa

Towards Universal Khmer Text Recognition

Khmer is a low-resource language characterized by a complex script, presenting significant challenge

BankNote-Net: Open dataset for assistive universal currency recognition

Millions of people around the world have low or no vision. Assistive software applications have been

Using articulatory feature detectors in progressive networks for multilingual low-resource phone recognition

Systems inspired by progressive neural networks, transferring information from end-to-end articulato