Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

masakhane-io/masakhane-news

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
mas
Hôte:
MasakhaNEWS: News Topic Classification for African Languages # MasakhaNEWS: News Topic Classification for African Languages paper | dataset >African languages are severely under-represented in NLP research due to lack of datasets covering several NLP tasks. While there are individual language specific datasets that are being expanded to different tasks, only a handful of NLP tasks (e.g. named entity recognition and machine translation) have standardized benchmark datasets covering several geographical and typologically-diverse African languages. In this paper, we develop MasakhaNEWS -- a new benchmark dataset for news topic classification covering 16 languages widely spoken in Africa. We provide an evaluation of baseline models by training classical machine learning models and fine-tuning several language models. Furthermore, we explore several alternatives to full fine-tuning of language models that are better suited for zero-shot and few-shot learning such as cross-lingual parameter-efficient fine-tuning (like MAD-X), pattern exploiting training (PET), prompting language models (like ChatGPT), and prompt-free sentence transformer fine-tuning (SetFit and Cohere Embedding API). Our evaluation in zero-shot setting shows the potential of prompting ChatGPT for news topic classification in low-resource African languages, achieving an average performance of 70 F1 points without leveraging additional supervision like MAD-X. In few-shot setting, we show that with as little as 10 examples per label, we achieved more than 90\% (i.e. 86.0 F1 points) of the performance of full supervised training (92.6 F1 points) leveraging the PET approach. ## Languages There are 16 languages available : - Amharic (amh) - English (eng) - French (fra) - Hausa (hau) - Igbo (ibo) - Lingala (lin) - Luganda (lug) - Oromo (orm) - Nigerian Pidgin (pcm) - Rundi (run) - chiShona (sna) - Somali (som) - Kiswahili (swą) - Tigrinya (tir) - isiXhosa (xho) - Yorùbá (yor) ### Monolingual Corpus We obtained the data mostly from BBC news, and a few other websites …

Visit

github.com

Tasks

news classificationtext classificationtopic classification

Languages

AmharicGandaHausaLingalaOromoOromo, Borana-Arsi-GujiRundiShonaSomaliSwahili+5

Similaires

masakhane-io/masakhane-mtmasakhane-io/masakhane-posmasakhane-io/masakhane-communitymasakhane-io/masakhane-khoekhoegowabmasakhane-io/masakhane-wazobia-datasetmasakhane-io/masakhanePreprocessor

masakhane-io/masakhane-mt

Machine Translation for Africa # Masakhane - A living collection of NLP projects for Africans, by A

masakhane-io/masakhane-pos

POS for African languages MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African Lang

masakhane-io/masakhane-community

All our community docs! Start here! Lets put Africa on the NLP Map # Masakhane - A living collectio

masakhane-io/masakhane-khoekhoegowab

Our collection of corpora for damara language # masakhane-khoekhoegowab Our collection of corpora f

masakhane-io/masakhane-wazobia-dataset

Some Nigerian Parallel Corpora: Yoroba, Igbo, Hausa, Urhobo

masakhane-io/masakhanePreprocessor

Building an effective preprocessing tool for African languages # `masakhanePreprocessor` An effecti