Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

An LLM-Driven Framework for Addressing Code-Switching and Orthographic Variance in Nigerian Language Data Collection

Domain:

natural language processing

Record type:

paper
Creator:
RosAghKai
Publisher:
Fed
Host:
Natural Language Processing (NLP) systems for low-resource languages continue to face significant challenges due to the poor quality and structural inconsistency of available textual data. In the Nigerian linguistic context, informal digital communication is characterized by frequent code-switching between English, Nigerian Pidgin, and indigenous languages, as well as high levels of orthographic variance. These characteristics reduce the effectiveness of conventional preprocessing pipelines, which are typically designed for monolingual and standardized text. Existing approaches based on rule-based filtering, statistical methods, or conventional language identification models often fail to accurately interpret multilingual and inconsistent language patterns. This paper presents a reasoning-driven preprocessing framework for improving Nigerian language data collection for NLP systems. The proposed approach leverages Large Language Models (LLMs) to perform context-aware language identification, relevance filtering, and orthographic normalization. Unlike traditional preprocessing methods that rely primarily on lexical or statistical matching, the framework treats language identification and normalization as contextual reasoning tasks. The system is designed to preserve semantic meaning while reducing spelling variability and dataset noise. To evaluate the framework, textual data was collected from informal online sources, including blogs, forums, and social media platforms containing multilingual Nigerian language content. Experimental results demonstrate that the proposed approach achieved an 85% reduction in noisy data while improving the handling of code-switched and orthographically inconsistent text. The findings highlight the potential of LLM-based preprocessing for enhancing data quality in low-resource NLP environments and provide a scalable foundation for developing more robust and inclusive language technologies for Nigerian languages.

Visit

doi.org

Tasks

code switchinglanguage identificationtext normalization

Licenses

https://creativecommons.org/licenses/by/4.0

Similar

KayusQoS: A Mobile-Driven Software Framework for Real-Time QoS Data Collection and Knowledge Discovery in Nigerian GSM NetworksAddressing Code-Switching in French/Algerian Arabic SpeechCode-switching in Nigerian hip-hop lyricsEducational language policy in an African country: Making a place for code-switching/translanguagingCorpus of English and Nigerian Pidgin Code-switching (CENCOS)Variation and language engineering in Yoruba-English code-switching

KayusQoS: A Mobile-Driven Software Framework for Real-Time QoS Data Collection and Knowledge Discovery in Nigerian GSM Networks

The Nigerian telecommunications industry relies largely on expensive and manually intensive drive-te

Addressing Code-Switching in French/Algerian Arabic Speech

International audience This study focuses on code-switching (CS) in French/Algerian A

Code-switching in Nigerian hip-hop lyrics

Educational language policy in an African country: Making a place for code-switching/translanguaging

Abstract In Ghana, plurilingual language use is the norm rather than the exception. It follows that

Corpus of English and Nigerian Pidgin Code-switching (CENCOS)

This dataset was compiled from a fieldwork in Nigeria in 2019. It features naturally occurring spoke

Variation and language engineering in Yoruba-English code-switching

This study deals with the identification and characterization of the variable features of code-switc