Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Harnessing Ensemble and Transformers for Sentiment Analysis and Emotion Detection in Hausa Text

Domain:

natural language processing

Record type:

paperdataset
Creator:
TaiKab
Publisher:
RSI
Host:
Understanding emotional tone and sentiment in text has driven significant advancements in Natural Language Processing (NLP), particularly in sentiment analysis and emotion detection. This study addresses the challenge of developing effective NLP tools for low-resource languages, focusing on the Hausa language. By leveraging ensemble methods and pre-trained transformer models like BERT and XLM-R, along with traditional classifiers such as Logistic Regression, SVM, Naive Bayes, Random Forest, and XGBoost, we aim to improve sentiment analysis and emotion detection for Hausa text. Utilizing a balanced sentiment dataset (9,958 samples) and a complex multi-label emotion dataset (19,757 samples across 11 categories), we benchmark individual classifiers, voting ensembles, and deep contextual models. For sentiment analysis, a Hard Voting Ensemble of TF-IDF-vectorized base learners achieved a highly competitive F1-score of 0.8748. However, Transformer models significantly outperformed traditional baselines, with Multilingual BERT (mBERT) achieving a peak F1-score of 0.8983. In the multi-label emotion detection task, individual traditional models struggled with label sparsity, yielding low Subset Accuracy scores (2.88% to 8.30%) and moderate Micro-F1 scores. Standard Hard Voting ensembles further underperformed due to discrete prediction conflicts. To resolve this, a Probability-based Majority Voting mechanism with calibrated thresholding (0.3) was introduced, boosting the Micro-F1 to 0.3825 and reducing the Hamming Loss to 0.1967. Ultimately, XLM-RoBERTa emerged as the superior architecture, achieving a Subset Accuracy of 0.1545, a Micro-F1 of 0.4275, and the lowest Hamming Loss of 0.1804. This research establishes a rigorous benchmark for Hausa NLP, highlighting the indispensable role of subword tokenization, contextual embeddings, and threshold calibration in handling the morphological richness and multi-label complexities of low-resource languages

Visit

doi.org

Tasks

emotion identificationsentiment analysistext classification

Languages

Hausa

Similar

Fine-tuning Multilingual Transformers for Hausa-English Sentiment AnalysisHausaNLP at SemEval-2025 Task 11: Hausa Text Emotion DetectionDz-Emotion: An Algerian Dialect Dataset for Text-Based Emotion DetectionUzEDSA: Uzbek Emotion and Sentiment Analysis DatasetComprehensive Hausa Language Processing Models for Text Summarization, Sentiment Analysis, Machine Translation, and Question AnsweringHarnessing Transformers for Enhancing Arabic Educational Assessment

Fine-tuning Multilingual Transformers for Hausa-English Sentiment Analysis

HausaNLP at SemEval-2025 Task 11: Hausa Text Emotion Detection

This paper presents our approach to multi-label emotion detection in Hausa, a low-resource African l

Dz-Emotion: An Algerian Dialect Dataset for Text-Based Emotion Detection

UzEDSA: Uzbek Emotion and Sentiment Analysis Dataset

UzEDSA (Uzbek Emotion and Sentiment Analysis Dataset) is an open-access dataset designed for emotion

Comprehensive Hausa Language Processing Models for Text Summarization, Sentiment Analysis, Machine Translation, and Question Answering

Harnessing Transformers for Enhancing Arabic Educational Assessment

This study introduces an intelligent scoring approach that leverages natural language processing and