Logo Lanfrica

AymanMansour/New-Lisan-Sudanese-TTS-Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Aym
Hôte:
Lisan Sudanese TTS Dataset A synthetic Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) dataset specifically for Sudanese Arabic. 1,878 high-quality sentences featuring 20 synthetic speakers (10 male, 10 female). Reconstructed from the Lisan-Sudanese Morphological Dataset (52K manually annotated social media tokens from Facebook/X). Only sentences with a diacritic density of >=25% were kept to ensure enough phonetic information for accurate synthesis. model: