Logo Lanfrica

Foundations and Frontiers of Swahili Natural Language Processing

Domaine:

natural language processing

Type de record:

paper
Créateur:
BenJar
Éditeur:
Afr
Hôte:
Swahili serves as a lingua franca for over 200 million speakers across East and Central Africa, functioning as a critical engine for digital societal interaction and economic integration. Despite its extensive demographic prominence, it remains acutely under-resourced within the computational domain of Natural Language Processing (NLP). This article presents a mathematically rigorous, systematic review of the Swahili NLP landscape, charting its historical trajectory from acute data scarcity to an emerging ecosystem of statistically specialised tools and representations. We comprehensively catalogue and mathematically categorise primary labelled datasets deployed for classification and generative tasks, including sentiment analysis, automatic speech recognition (ASR), and abstractive summarisation. Furthermore, we analyse the parameter optimisation, tokenisation dynamics, and architectural efficacy of dedicated foundational language models—such as SwahBERT, AfriBERTa, and specialised Large Language Models (LLMs)—mapping their performance bounds against Swahili’s highly agglutinative morphology. Beyond a mere structural inventory, this review critically exposes the “resource quality gap”, identifying profound statistical deficiencies in domain-specific corpora across high-stakes fields such as healthcare and law, whilst highlighting an acute lack of dialectal variance and fine-grained annotation schemes. The analysis concludes that while generalised parameter spaces have expanded, the absence of unified benchmarking infrastructures impedes systematic algorithmic validation. Finally, we propose actionable, statistically sound methodologies for future research, advocating for the collaborative construction of comprehensive multi-task benchmarks to bridge the digital divide for the Swahili-speaking populace.

Languages

Similaires