Logo Lanfrica

neh7777/Pretraining-V1

Domaine:

natural language processing

Type de record:

dataset
Créateur:
neh
Hôte:
A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio.