Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Ibom NLP: A Step Toward Inclusive Natural Language Processing for Nigeria's Minority Languages

Domain:

natural language processing

Record type:

datasetpaper

Nigeria is the most populous country in Africa with a population of more than 200 million people. More than 500 languages are spoken in Nigeria and it is one of the most linguistically diverse countries in the world. Despite this, natural language processing (NLP) research has mostly focused on the following four languages: Hausa, Igbo, Nigerian-Pidgin, and Yoruba (i.e <1% of the languages spoken in Nigeria). This is in part due to the unavailability of textual data in these languages to train and apply NLP algorithms. In this work, we introduce ibom -- a dataset for machine translation and topic classification in four Coastal Nigerian languages from the Akwa Ibom State region: Anaang, Efik, Ibibio, and Oro. These languages are not represented in Google Translate or in major benchmarks such as Flores-200 or SIB-200. We focus on extending Flores-200 benchmark to these languages, and further align the translated texts with topic labels based on SIB-200 classification dataset. Our evaluation shows that current LLMs perform poorly on machine translation for these languages in both zero-and-few shot settings. However, we find the few-shot samples to steadily improve topic classification with more shots.

Visit

arxiv.orgview datasetcollection on huggingface

Tasks

machine translationtopic classificationtext classification

Languages

AnaangEfikIbibioOro

Similar

Natural language processing for African languages NLP for African languagesNalediMsiya/Natural-Language-Processing-NLP-Natural Language Processing (NLP) tools - multilingual and low-resource languagesNatural language processing for African languagesNatural Language Processing (NLP) for Requirements Engineering: A Systematic Mapping Study DatasetNatural Language Processing (NLP) Techniques for Afan Oromo Text Analysis

Natural language processing for African languages NLP for African languages

NalediMsiya/Natural-Language-Processing-NLP-

This is a dataset of wars recorded in Africa from the 1990's to 2023 , NLP is applied to analyze tex

Natural Language Processing (NLP) tools - multilingual and low-resource languages

Natural language processing for African languages

Recent advances in pre-training of word embeddings and language models leverage large amounts of unlabelled texts and self-supervised learning to learn distributed representations that have significantly improved the performance of deep learning models on a large v

Natural Language Processing (NLP) for Requirements Engineering: A Systematic Mapping Study Dataset

We conducted a systematic mapping study to address and survey the issues of using Natural Language P

Natural Language Processing (NLP) Techniques for Afan Oromo Text Analysis

Natural Language Processing (NLP) has emerged as a transformative tool for analyzing and understandi