Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

BBC Igbo-Pidgin Gold-Standard NLP Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Bytte AI

This corpus contains high-quality BBC Igbo and BBC Pidgin news snippets annotated by Bytte AI for intent, sentiment, content quality, sentence segmentation, and Named Entity Recognition (NER). Comprising 63 Igbo and 91 Pidgin samples per task, it provides a rich, multilingual benchmark for NLP research. Designed for low-resource African languages, it enables model training, evaluation, and benchmarking on real-world journalistic text.

Visit

figshare.comgithub.comhuggingface.co

Tasks

named entity recognitionsentiment analysissentence segmentationinformation extractiontext classification

Languages

IgboPidgin, Nigerian

Tags

African datasetsIgbo datasetsPidgin DatasetsNigeriaWest AfricaEditorialBBC

Licenses

CC-BY-4.0

Similar

<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>bytte-ai/igbo-pidgin-nlp-corpus

<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>

This corpus is a high-quality, manually annotated collection of BBC Igbo and BBC Pidgin

bytte-ai/igbo-pidgin-nlp-corpus

Sample dataset: High-quality annotated data for Nigerian Igbo and Pidgin English NLP 🤗 Hugging Face