Logo Lanfrica

{language_name} Speech-Text Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
mic
Host:
This dataset is a collection of aligned audio-text pairs in Lomwe, extracted from the CMU Wilderness dataset. It is useful for tasks such as: Speech recognition (ASR) Text-to-speech (TTS) Language modeling for low-resource languages Each entry in the dataset contains: audio: A .wav file sampled at 16kHz text: A transcription of the spoken audio in Lomwe (digits removed) audio