Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

POWSM: A Phonetic Open Whisper-Style Speech Foundation Model

Domain:

natural language processing

Record type:

papermodeldatasetsoftware
Creator:
Li,ChaBhaYeo
Host:avatar
Recent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-phoneme conversion (G2P), and phoneme-to-grapheme conversion (P2G). Despite their conceptual similarity, these tasks have largely been studied in isolation, each relying on task-specific architectures and datasets. In this paper, we introduce POWSM (Phonetic Open Whisper-style Speech Model), the first unified framework capable of jointly performing multiple phone-related tasks. POWSM enables seamless conversion between audio, text (graphemes), and phones, opening up new possibilities for universal and low-resource speech processing. Our model outperforms or matches specialized PR models of similar size (Wav2Vec2Phoneme and ZIPA) while jointly supporting G2P, P2G, and ASR. Our training data, code and models are released to foster open science. 18 pages, under review. Model available at huggingface.co

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageAudio and Speech Processing

Similar

BULaMU: An Open Foundation Model for LugandaLanguages in Whisper-Style Speech Encoders Align Both Phonetically and Semanticallynelovoice/nelo-whisper-darja-speech-modelTowards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource LanguageNyengor/open-foundation-west-africa-projectNeural Networks For Speech Recognition Of A Phonetic Language

BULaMU: An Open Foundation Model for Luganda

Uganda, colloquially referred to as the “pearl of Africa”, is home to

Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically

Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. S

nelovoice/nelo-whisper-darja-speech-model

# nelo-whisper-darja-speech-model Fine-tuned Whisper model for Moroccan Darija speech recognition.

Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language

The increase in technological adoption worldwide comes with demands for novel tools to be used by th

Nyengor/open-foundation-west-africa-project

# Galamsay Data Analysis & REST API This project analyzes illegal small-scale mining (Galamsey) act

Neural Networks For Speech Recognition Of A Phonetic Language

The goal of this thesis is to explore a possibility for a viable alternative/replacement to the Amha