This dataset comprises audio recordings of Nigerian Pidgin English speech aligned with textual transcriptions. The dataset is structured into 16 folders, each containing audio files and a corresponding audio-text mapping file.
The audio clips are short, typically ranging from 1 to 38 seconds, and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file.
The textual content used in this dataset originates from a variety of written and spoken sources in Nigerian Pidgin English, including narrative texts, conversational exchanges, news-style content, and everyday speech samples. These texts were segmented into short utterances suitable for read speech and TTS modelling.