Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

An LLM-Driven Framework for Addressing Code-Switching and Orthographic Variance in Nigerian Language Data Collection

Domaine:

natural language processing

Type de record:

paper
Créateur:
RosAghKai
Éditeur:
Fed
Hôte:
Natural Language Processing (NLP) systems for low-resource languages continue to face significant challenges due to the poor quality and structural inconsistency of available textual data. In the Nigerian linguistic context, informal digital communication is characterized by frequent code-switching between English, Nigerian Pidgin, and indigenous languages, as well as high levels of orthographic variance. These characteristics reduce the effectiveness of conventional preprocessing pipelines, which are typically designed for monolingual and standardized text. Existing approaches based on rule-based filtering, statistical methods, or conventional language identification models often fail to accurately interpret multilingual and inconsistent language patterns. This paper presents a reasoning-driven preprocessing framework for improving Nigerian language data collection for NLP systems. The proposed approach leverages Large Language Models (LLMs) to perform context-aware language identification, relevance filtering, and orthographic normalization. Unlike traditional preprocessing methods that rely primarily on lexical or statistical matching, the framework treats language identification and normalization as contextual reasoning tasks. The system is designed to preserve semantic meaning while reducing spelling variability and dataset noise. To evaluate the framework, textual data was collected from informal online sources, including blogs, forums, and social media platforms containing multilingual Nigerian language content. Experimental results demonstrate that the proposed approach achieved an 85% reduction in noisy data while improving the handling of code-switched and orthographically inconsistent text. The findings highlight the potential of LLM-based preprocessing for enhancing data quality in low-resource NLP environments and provide a scalable foundation for developing more robust and inclusive language technologies for Nigerian languages.

Visit

doi.org

Tasks

code switchinglanguage identificationtext normalization

Licenses

https://creativecommons.org/licenses/by/4.0