Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African Language

Domain:

natural language processing

Record type:

paperdataset
Creator:
ZanAbdAdoBic
Host:avatar
The development of Natural Language Processing (NLP) tools for low-resource languages is critically hindered by the scarcity of annotated datasets. This paper addresses this fundamental challenge by introducing HausaMovieReview, a novel benchmark dataset comprising 5,000 YouTube comments in Hausa and code-switched English. The dataset was meticulously annotated by three independent annotators, demonstrating a robust agreement with a Fleiss' Kappa score of 0.85 between annotators. We used this dataset to conduct a comparative analysis of classical models (Logistic Regression, Decision Tree, K-Nearest Neighbors) and fine-tuned transformer models (BERT and RoBERTa). Our results reveal a key finding: the Decision Tree classifier, with an accuracy and F1-score 89.72% and 89.60% respectively, significantly outperformed the deep learning models. Our findings also provide a robust baseline, demonstrating that effective feature engineering can enable classical models to achieve state-of-the-art performance in low-resource contexts, thereby laying a solid foundation for future research. Keywords: Hausa, Kannywood, Low-Resource Languages, NLP, Sentiment Analysis Masters Thesis, a Dataset Paper

Visit

arxiv.org

Tasks

sentiment analysistext classification

Languages

Hausa

Tags

Computation and LanguageArtificial Intelligence

Similar

HausaMovieReview: A Manually Annotated Sentiment Dataset for Hausa Language NLPSentiMaithili: A Benchmark Dataset for Sentiment and Reason Generation for the Low-Resource Maithili LanguageTunDC: a public benchmark dataset for sentiment analysis and language modeling in the Tunisian dialectAfriSenti: A Twitter Sentiment Analysis Benchmark for African LanguagesIammteo/Sentiment-analysis-model-for-Nigerian-pidgin-Low-resource-language-Enhancing Sentiment Analysis in Amharic: Leveraging Transformer-Based Language Model for Low-Resource African Languages

HausaMovieReview: A Manually Annotated Sentiment Dataset for Hausa Language NLP

PAIDeF SuperAI 2025 Conference

HausaMovieReview: A Manually Annotated Se

SentiMaithili: A Benchmark Dataset for Sentiment and Reason Generation for the Low-Resource Maithili Language

Developing benchmark datasets for low-resource languages poses significant challenges, primarily due

TunDC: a public benchmark dataset for sentiment analysis and language modeling in the Tunisian dialect

The development of natural language processing (NLP) applications has increasingly focused on dialec

AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages

Africa is home to over 2000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages

Iammteo/Sentiment-analysis-model-for-Nigerian-pidgin-Low-resource-language-

This is a machine learning project of a Nigerian Pidgin sentiment analysis model specifically design

Enhancing Sentiment Analysis in Amharic: Leveraging Transformer-Based Language Model for Low-Resource African Languages