Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

Domaine:

natural language processing

Type de record:

paperdatasetmodelsoftware
Créateur:
GuzBeyAlaJeo
Hôte:avatar
Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially improved synthetic speech quality, yet these gains remain unevenly distributed across the world's languages. Existing models are still dominated by a small set of high-resource languages, while many studies of low-resource TTS are simulated on artificially downsampled high-resource corpora that do not reflect the orthographic variation and limited phonetic coverage encountered in genuinely underrepresented settings. As such, we introduce OpenBibleTTS, which is a large-scale benchmark for low-resource speech synthesis spanning 37 underrepresented languages. Moreover, a systematic comparison of various TTS architectures and large-scale speech generation models is conducted across in-domain Biblical text and out-of-domain material. Results show that no single system dominates across languages and metrics: Gemini-TTS achieves the highest listener ratings on most evaluated languages, but monolingual EveryVoice models trained on OpenBibleTTS remain strongest for intelligibility and are preferred in several African languages, while open from-scratch systems degrade sharply on out-of-domain text, revealing a persistent gap between broad multilingual coverage and reliable synthesis quality in underserved linguistic communities. We complement automatic evaluation with subjective human judgments, and open-source all processed datasets, alignments, and trained models to support future low-resource TTS research.

Visit

arxiv.org

Tasks

speech processingtext to speech

Tags

Computation and LanguageSound

Similaires

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource LanguagesLarge Multimodal Models for Low-Resource Languages: A SurveyAdaptive and Efficient Large Language Models for Low-Resource African LanguagesLarge Language Models Adaptation for Low-resource Languages: The Case for African LanguagesVisually Grounded Speech Models for Low-resource Languages and Cognitive ModellingContinual-learning for Modelling Low-Resource Languages from Large Language Models

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

Speech large language models (SLLMs) built on speech encoders, adapters, and LLMs demonstrate remark

Large Multimodal Models for Low-Resource Languages: A Survey

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) fo

Adaptive and Efficient Large Language Models for Low-Resource African Languages

PAIDeF SuperAI 2025 Conference

Adaptive and Efficient Large Language Mod

Large Language Models Adaptation for Low-resource Languages: The Case for African Languages

David Ifeoluwa Adelani (Supervisor) Despite remarkable advances in Large Language Models (LLMs), Afr

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech p

Continual-learning for Modelling Low-Resource Languages from Large Language Models

Modelling a language model for a multi-lingual scenario includes several potential challenges, among