Eng-PidginEdu is an English–Nigerian Pidgin parallel corpus developed to support machine translation, multilingual NLP, and educational accessibility research for low-resource African languages.
The dataset contains 26,232 parallel sentence pairs collected across 8 Nigerian secondary school subjects, comprising approximately 1.09 million total tokens with high-quality sentence-level alignment and human-validated translations.