
Eng-PidginBioData is a domain-specific parallel corpus for English ↔ Nigerian Pidgin machine translation focused on biological and scientific texts. The dataset contains 2,300 sentence pairs extracted from open-source biological research papers and manually translated into Nigerian Pidgin.
The dataset was created to support research in low-resource language machine translation, particularly for Nigerian Pidgin, which despite being spoken by millions of people remains underrepresented in scientific literature and NLP resources.