Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

News Classification in Low‐Resource Languages: Insights From Transformer and Baseline Models

Domaine:

natural language processing

Type de record:

paper
Créateur:
WubAli
Éditeur:
WILEY
Hôte:
ABSTRACT News text classification in low‐resource languages such as Amharic is challenging due to limited annotated data, rich morphology, class imbalance, and strong semantic overlaps (SOs) among news categories. Addressing these challenges is critical for reliable information organization in socially important domains, including media, education, policymaking, and so forth. In this work, we propose a unified framework for Amharic multi‐class news classification that integrates two state‐of‐the‐art (SOTA) Transformer‐based models with two traditional machine learning approaches under identical experimental setups. Specifically, we fine‐tune AfriBERTa, a monolingual RoBERTa‐based model pre‐trained on African languages, and AfroXLMR, a multilingual variant of XLM‐R, and compare them with TF‐IDF combined with logistic regression and Word2Vec combined with a multi‐layer perceptron. To the best of our knowledge, this is the first study to jointly benchmark these Transformer‐based and traditional models on an Amharic news dataset using stratified five‐fold cross‐validation (CV) and class‐balanced training. To enhance interpretability and real‐world usability, we deploy a real‐time Gradio‐based graphical user interface that exposes class‐wise probability distributions, enabling transparent analysis of SO across classes. The experimental results on a hold‐out test set show that AfriBERTa achieves the best performance, with a macro F 1 score of 94.12%, followed by AfroXLMR with 92.42%, while TF‐IDF + LR and Word2Vec + MLP achieve macro F 1 scores of 90.34% and 88.17%, respectively. All results are validated through statistical significance testing and comparative evaluation against zero‐shot large language models (LLMs), including ChatGPT‐4o, Gemini Pro, and Claude 3, where the proposed models consistently outperform due to language‐specific adaptability. Macro F 1 is used as the primary evaluation metric to ensure fair assessment under class imbalance and SO. Overall, this work provides a reproducible and interpretable benchmark for low‐resource news classification and contributes to research in explainable artificial intelligence and African language processing, with future directions including multimodal news text classifications.

Visit

doi.org

Tasks

news classificationtext classificationtopic classification

Languages

Amharic

Licenses

http://onlinelibrary.wiley.com/termsAndConditions#vorhttp://doi.wiley.com/10.1002/tdm_license_1.1