Logo Lanfrica

Hypa Speech 10k

Domaine:

natural language processing

Type de record:

dataset
Créateur:
hyp
Hôte:
A multilingual instruction-tuning dataset covering translation,transcription, and language detection. Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets.