Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Who Wrote This? Identifying Machine vs Human-Generated Text in Hausa

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
SanSoyImaMus
Hôte:avatar
The advancement of large language models (LLMs) has allowed them to be proficient in various tasks, including content generation. However, their unregulated usage can lead to malicious activities such as plagiarism and generating and spreading fake news, especially for low-resource languages. Most existing machine-generated text detectors are trained on high-resource languages like English, French, etc. In this study, we developed the first large-scale detector that can distinguish between human- and machine-generated content in Hausa. We scrapped seven Hausa-language media outlets for the human-generated text and the Gemini-2.0 flash model to automatically generate the corresponding Hausa-language articles based on the human-generated article headlines. We fine-tuned four pre-trained Afri-centric models (AfriTeVa, AfriBERTa, AfroXLMR, and AfroXLMR-76L) on the resulting dataset and assessed their performance using accuracy and F1-score metrics. AfroXLMR achieved the highest performance with an accuracy of 99.23% and an F1 score of 99.21%, demonstrating its effectiveness for Hausa text detection. Our dataset is made publicly available to enable further research.

Visit

arxiv.org

Tasks

text classification

Languages

Hausa

Tags

Computation and Language

Similaires

shizaamir1615/human-vs-machine-text-detectorIdentifying Text Classification Failures in Multilingual AI-Generated ContentDual-BERT Adversarial Model for Text Normalization in Hausa User-Generated ContentsYabsera-Haile/Human-vs-Machine-Translation-Detection-CodebaseDetecting Machine-Generated Arabic Text: AraBERT–LSTM for Trustworthy Low-Resource NLPIdentifying Sentiments in Algerian Code-switched User-generated Comments

shizaamir1615/human-vs-machine-text-detector

Group 23- Detecting machine generated content in low resource African languages Human vs Machine Te

Identifying Text Classification Failures in Multilingual AI-Generated Content

With the rising popularity of generative AI tools, the nature of apparent classification failures by

Dual-BERT Adversarial Model for Text Normalization in Hausa User-Generated Contents

Abstract This paper presents an innovative Dual-BERT Generative Adversarial Networ

Yabsera-Haile/Human-vs-Machine-Translation-Detection-Codebase

This repository is used for reproducibility of a study of human vs machine translation detection acr

Detecting Machine-Generated Arabic Text: AraBERT–LSTM for Trustworthy Low-Resource NLP

Deepfake text generation has emerged as a serious challenge in the age of advanced language models,

Identifying Sentiments in Algerian Code-switched User-generated Comments

We present in this paper our work on Algerian language, an under-resourced North African colloquial Arabic variety, for which we built a comparably large corpus of more than 36,000 code-switched user-generated comments annotated for sentiments. We opted for this da