Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

POWSM: A Phonetic Open Whisper-Style Speech Foundation Model

Domaine:

natural language processing

Type de record:

papermodeldatasetsoftware
Créateur:
Li,ChaBhaYeo
Hôte:avatar
Recent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-phoneme conversion (G2P), and phoneme-to-grapheme conversion (P2G). Despite their conceptual similarity, these tasks have largely been studied in isolation, each relying on task-specific architectures and datasets. In this paper, we introduce POWSM (Phonetic Open Whisper-style Speech Model), the first unified framework capable of jointly performing multiple phone-related tasks. POWSM enables seamless conversion between audio, text (graphemes), and phones, opening up new possibilities for universal and low-resource speech processing. Our model outperforms or matches specialized PR models of similar size (Wav2Vec2Phoneme and ZIPA) while jointly supporting G2P, P2G, and ASR. Our training data, code and models are released to foster open science. 18 pages, under review. Model available at huggingface.co

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageAudio and Speech Processing

Similaires

BULaMU: An Open Foundation Model for LugandaLanguages in Whisper-Style Speech Encoders Align Both Phonetically and Semanticallynelovoice/nelo-whisper-darja-speech-modelTowards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource LanguageNyengor/open-foundation-west-africa-projectNeural Networks For Speech Recognition Of A Phonetic Language

BULaMU: An Open Foundation Model for Luganda

Uganda, colloquially referred to as the “pearl of Africa”, is home to

Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically

Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. S

nelovoice/nelo-whisper-darja-speech-model

# nelo-whisper-darja-speech-model Fine-tuned Whisper model for Moroccan Darija speech recognition.

Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language

The increase in technological adoption worldwide comes with demands for novel tools to be used by th

Nyengor/open-foundation-west-africa-project

# Galamsay Data Analysis & REST API This project analyzes illegal small-scale mining (Galamsey) act

Neural Networks For Speech Recognition Of A Phonetic Language

The goal of this thesis is to explore a possibility for a viable alternative/replacement to the Amha