Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
KhaAli
Hôte:avatar
Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic -- the most widely understood Arabic dialect -- severely under-resourced. We address this gap by introducing NileTTS: 38 hours of transcribed speech from two speakers across diverse domains including medical, sales, and general conversations. We construct this dataset using a novel synthetic pipeline: large language models (LLM) generate Egyptian Arabic content, which is then converted to natural speech using audio synthesis tools, followed by automatic transcription and speaker diarization with manual quality verification. We fine-tune XTTS v2, a state-of-the-art multilingual TTS model, on our dataset and evaluate against the baseline model trained on other Arabic dialects. Our contributions include: (1) the first publicly available Egyptian Arabic TTS dataset, (2) a reproducible synthetic data generation pipeline for dialectal TTS, and (3) an open-source fine-tuned model. All resources are released to advance Egyptian Arabic speech synthesis research. 8 pages, 2 figures, EACL26

Visit

arxiv.org

Tasks

text to speechspeech processing

Tags

Computation and Language

Similaires

sayleee1/speech-to-text-pipelineEnhancing Crowdsourced Audio for Text-to-Speech ModelsText-To-Speech Data Augmentation for Low Resource Speech RecognitionYoruba TTS (text-to-speech) training datasetBuilding Text-to-Speech Models for Low-Resourced Languages from Crowdsourced Datatheonlyamos/ghana-languages-speech-to-speech-pipeline

sayleee1/speech-to-text-pipeline

Web app to collect Amharic speech data for NLP model training # 🗣️ Speech-to-Text Data Collection P

Enhancing Crowdsourced Audio for Text-to-Speech Models

High-quality audio data is a critical prerequisite for training robust text-to-speech models, which

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Nowadays, the main problem of deep learning techniques used in the development of automatic speech r

Yoruba TTS (text-to-speech) training dataset

Textbook audio archive size: total 36M archive created: 8 July 2011 mp3 file size ======== ==== 01-

Building Text-to-Speech Models for Low-Resourced Languages from Crowdsourced Data

Text-to-speech (TTS) models have expanded the scope of digital inclusivity by becoming a basis for a

theonlyamos/ghana-languages-speech-to-speech-pipeline

A comprehensive pipeline for building multilingual Speech-to-Speech (S2S) AI systems for Ghanaian la