Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech

Domain:

natural language processing

Record type:

paperdataset
Creator:
Rah
Host:avatar
Text-to-speech (TTS) evaluation for low-resource non-Latin-script languages can fail when it relies on a single ASR round-trip word error rate (WER). A system may produce no audio, speak a neighbouring language, preserve target script text only in an ASR transcript, or sound unnatural to native listeners. We introduce INSV (Intelligibility, Naturalness, Script fidelity, and Verification), a reporting framework that separates these cases. This paper reports INSV-A, the automated screening subset: synthesis completion, ASR WER/CER, transcript Script Fidelity Rate, and audio language identification. Native MOS and phonetic annotation are specified but not claimed in this release. We instantiate INSV-A as PashtoTTS-Bench, a dated benchmark for Pashto TTS. The April-May 2026 run evaluates Edge GulNawaz, Edge Latifa, OmniVoice clone, OmniVoice auto, and an Urdu negative control on 200 FLEURS and 200 filtered Common Voice 24 prompts. Under the independent omniASR_CTC_300M_v2, OmniVoice auto has the lowest WER (24.1% FLEURS, 27.4% CV24), followed by Edge GulNawaz (32.8%, 39.5%), Edge Latifa (35.6%, 47.7%), and OmniVoice clone (45.4%, 34.8%). WER below the natural-speech baseline reflects clean synthetic audio and should not be read as better than native speech. Whisper Large V3 returns 0.0% Pashto labels on checked Pashto TTS audio, while MMS-LID-4017 and SpeechBrain VoxLingua107 separate Pashto outputs from the Urdu control. The release provides provider metadata, per-sentence scores, LID audits, failure logs, and scripts for adding systems.

Visit

arxiv.org

Tasks

text to speechspeech processing

Tags

Computation and LanguageSound

Similar

Script Transliteration Effects on Multilingual NER Robustness in Low-Resource Non-Latin LanguagesAutomatic Keyboard Layout Design for Low-Resource Latin-Script LanguagesAdversarial Text-to-Speech for low-resource languagesText-To-Speech Data Augmentation for Low Resource Speech RecognitionText to Speech System for Meitei Mayek Script Text-to-Speech Synthesis Using Found Data for Low-Resource Languages

Script Transliteration Effects on Multilingual NER Robustness in Low-Resource Non-Latin Languages

Pretrained multilingual language models have become a common tool in transferring NLP capabilities t

Automatic Keyboard Layout Design for Low-Resource Latin-Script Languages

We present our approach to automatically designing and implementing keyboard layouts on mobile devic

Adversarial Text-to-Speech for low-resource languages

Improving the adversarial TTS models for low-resource languages by utilizing the high-frequency similarities between the different languages.

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Nowadays, the main problem of deep learning techniques used in the development of automatic speech r

Text to Speech System for Meitei Mayek Script

This paper presents the development of a Text-to-Speech (TTS) system for the Manipuri language using

Text-to-Speech Synthesis Using Found Data for Low-Resource Languages

Text-to-speech synthesis is a key component of interactive, speech-based systems. Typically, buildi