Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Active Learning with Non-Uniform Costs for African Natural Language Processing

Domaine:

natural language processing

Type de record:

paper
Créateur:
AssAroCheDos
Éditeur:
Und
Hôte:avatar
Labeling datasets for African languages presents substantial challenges due to the diverse settings in which annotations are collected, resulting in highly variable labeling costs. These costs can vary with task complexity, annotator expertise, and data availability. Yet, most active learning (AL) frameworks assume uniform annotation costs, limiting their applicability in real-world, resource-constrained scenarios. To address this, we introduce KnapsackBALD, a novel cost-aware active learning method that integrates the BatchBALD acquisition strategy with a 0-1 Knapsack optimization objective to select informative and budget-efficient samples. We evaluate KnapsackBALD on the MasakhaNEWS dataset, a multilingual news classification benchmark covering 11 African languages. Our method consistently outperforms seven strong active learning baselines, including BALD, BatchBALD, and stochastic sampling variants such as PowerBALD and Softmax-BALD, across all three cost scenarios. The performance gap widens as annotation cost imbalances become more extreme, demonstrating the robustness of KnapsackBALD under practical constraints. These findings underscore the need for cost-sensitive acquisition in AL pipelines for African language NLP and beyond.

Visit

doi.orgunderline.io

Tasks

news classificationtext classificationtopic classification

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similaires

Natural language processing for African languagesNatural language processing for African languages NLP for African languagesTransfer Learning for Natural Language Processing by Paul AzunreKencorpus: Kenyan Languages Corpus for Machine Learning and Natural Language ProcessingKencorpus: Kenyan Languages Corpus for Natural Language Processing and Machine LearningReplication Data for Igbo Natural Language Processing Tasks I Igbo Synchronised Corpus for Natural Language Processing Tasks

Natural language processing for African languages

Recent advances in pre-training of word embeddings and language models leverage large amounts of unlabelled texts and self-supervised learning to learn distributed representations that have significantly improved the performance of deep learning models on a large v

Natural language processing for African languages NLP for African languages

Transfer Learning for Natural Language Processing by Paul Azunre

Training deep learning NLP models from scratch is costly, time-consuming, and requires massive amounts of data. In Transfer Learning for Natural Language Processing, DARPA researcher Paul Azunre reveals cutting-edge transfer learning techniques that apply customiza

Kencorpus: Kenyan Languages Corpus for Machine Learning and Natural Language Processing

This project collected text and speech corpora for three languages in Kenya: Kiswahili, Dholuo and 3 Luhya dialects (Lumarachi, Logooli and Lubukusu). Primary data was collected from the respective language communities, which included Indigenous stories and narrati

Kencorpus: Kenyan Languages Corpus for Natural Language Processing and Machine Learning

Kencorpus: Kenyan Languages Corpus for Natural Language Processing and Machine Learning

Poster presented at the Deep Learning Indaba 2022 by Lillian Wanzare

Replication Data for Igbo Natural Language Processing Tasks I Igbo Synchronised Corpus for Natural Language Processing Tasks

The Igbo synchronised corpus (IgboSynCorp) is an annotated corpus of spoken Igbo created by a team o