Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

HausaNLP: Current Status, Challenges and Future Directions for Hausa Natural Language Processing

Domaine:

natural language processing

Type de record:

paperprojectdataset
Créateur:
Muhammad, Shamsuddeen HassanAhmad, Ibrahim SaidAbdulmumin, IdrisLaw
Hôte:avatar
Hausa Natural Language Processing (NLP) has gained increasing attention in recent years, yet remains understudied as a low-resource language despite having over 120 million first-language (L1) and 80 million second-language (L2) speakers worldwide. While significant advances have been made in high-resource languages, Hausa NLP faces persistent challenges, including limited open-source datasets and inadequate model representation. This paper presents an overview of the current state of Hausa NLP, systematically examining existing resources, research contributions, and gaps across fundamental NLP tasks: text classification, machine translation, named entity recognition, speech recognition, and question answering. We introduce HausaNLP (catalog.hausanlp.org), a curated catalog that aggregates datasets, tools, and research works to enhance accessibility and drive further development. Furthermore, we discuss challenges in integrating Hausa into large language models (LLMs), addressing issues of suboptimal tokenization and dialectal variation. Finally, we propose strategic research directions emphasizing dataset expansion, improved language modeling approaches, and strengthened community collaboration to advance Hausa NLP. Our work provides both a foundation for accelerating Hausa NLP progress and valuable insights for broader multilingual NLP research.

Visit

arxiv.org

Languages

Hausa

Tags

Computation and Language

Similaires

Natural Language Processing for Tigrinya: Current State and Future DirectionsA Survey on Multilingual Natural Language Processing: Data, Models, Evaluation, and Future DirectionsNatural Language Processing in Ethiopian Languages: Current State, Challenges, and OpportunitiesExploring Gendered Challenges in Accessing Education for Sustainable Development in the Global South: Current Status and Future DirectionshauWE: Hausa Words Embedding for Natural Language ProcessingNatural language processing: Understanding the current landscape

Natural Language Processing for Tigrinya: Current State and Future Directions

Despite being spoken by millions of people, Tigrinya remains severely underrepresented in Natural La

A Survey on Multilingual Natural Language Processing: Data, Models, Evaluation, and Future Directions

While natural language processing (NLP) has advanced for major languages, most of the world’s 7,000

Natural Language Processing in Ethiopian Languages: Current State, Challenges, and Opportunities

This survey delves into the current state of natural language processing (NLP) for four Ethiopian la

Exploring Gendered Challenges in Accessing Education for Sustainable Development in the Global South: Current Status and Future Directions

Education for Sustainable Development (ESD) is a critical framework for addressing environmental deg

hauWE: Hausa Words Embedding for Natural Language Processing

Words embedding (distributed word vector representations) have become an essential component of many natural language processing (NLP) tasks such as machine translation, sentiment analysis, word analogy, named entity recognition and word similarity. Despite this, t

Natural language processing: Understanding the current landscape

We spoke with two researchers in Natural Language Processing (NLP) to understand their perspective o