A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM)
Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC).
The NAC Team processed the original long-form recordings by:
Segmenting the audio into sentence-level clips.
Transcribing/Aligning the text to create paired audio-text data suitable for ASR training.