Logo Lanfrica

Makhuwa Trigrams Speech-Text Parallel Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
mic
Hôte:
This dataset contains 154253 parallel speech-text pairs for Makhuwa, a language spoken primarily in Mozambique. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Language: Makhuwa - vmw