Logo Lanfrica

Eng-PidginBioData: English–Nigerian Pidgin Biology Translation Dataset

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Flora Oladipupo

Eng-PidginBioData is a domain-specific parallel corpus for English ↔ Nigerian Pidgin machine translation focused on biological and scientific texts. The dataset contains 2,300 sentence pairs extracted from open-source biological research papers and manually translated into Nigerian Pidgin.

The dataset was created to support research in low-resource language machine translation, particularly for Nigerian Pidgin, which despite being spoken by millions of people remains underrepresented in scientific literature and NLP resources.