Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BBC Igbo-Pidgin Gold-Standard NLP Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Bytte AI

This corpus contains high-quality BBC Igbo and BBC Pidgin news snippets annotated by Bytte AI for intent, sentiment, content quality, sentence segmentation, and Named Entity Recognition (NER). Comprising 63 Igbo and 91 Pidgin samples per task, it provides a rich, multilingual benchmark for NLP research. Designed for low-resource African languages, it enables model training, evaluation, and benchmarking on real-world journalistic text.

Visit

figshare.comgithub.comhuggingface.co

Tasks

named entity recognitionsentiment analysissentence segmentationinformation extractiontext classification

Languages

IgboPidgin, Nigerian

Tags

African datasetsIgbo datasetsPidgin DatasetsNigeriaWest AfricaEditorialBBC

Licenses

CC-BY-4.0

Similaires

<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>bytte-ai/igbo-pidgin-nlp-corpus

<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>

This corpus is a high-quality, manually annotated collection of BBC Igbo and BBC Pidgin

bytte-ai/igbo-pidgin-nlp-corpus

Sample dataset: High-quality annotated data for Nigerian Igbo and Pidgin English NLP 🤗 Hugging Face