A unified Nigerian Pidgin English speech-to-text dataset that combines
publicly available Pidgin ASR sources into a single train / validation /
test setup with a consistent schema. Built for fine-tuning Whisper-family
models on Nigerian Pidgin (Naija, pcm).
~8.6 hours, 4,278 clips, 10 source speakers, 16 kHz mono WAV.
Used to train michaelodafe/whisper-pidgin-v1
(21.37% WER on the test split, beating the published Wav2Vec2-XLSR-53