Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Task-Adaptive Pre-Training for Boosting Learning With Noisy Labels: A Study on Text Classification for African Languages

Domain:

natural language processing

Record type:

paper
For high-resource languages like English, text classification is a well-studied task. The performance of modern NLP models easily achieves an accuracy of more than 90\% in many standard datasets for text classification in English \citep{xie2019unsupervised, Yang2019, Zaheer2020}. However, text classification in low-resource languages is still challenging due to the lack of annotated data. Although methods like weak supervision and crowdsourcing can help ease the annotation bottleneck, the annotations obtained by these methods contain label noise. Models trained with label noise may not generalize well. To this end, a variety of noise-handling techniques have been proposed to alleviate the negative impact caused by the errors in the annotations (for extensive surveys see \citep{hedderich-etal-2021-survey, DBLP:journals/kbs/AlganU21}). In this work, we experiment with a group of standard noisy-handling methods on text classification tasks with noisy labels. We study both simulated noise and realistic noise induced by weak supervision. Moreover, we find task-adaptive pre-training techniques \citep{DBLP:conf/acl/GururanganMSLBD20} are beneficial for learning with noisy labels.

Visit

openreview.net

Tasks

text classification

Languages

HausaYoruba

Tags

africanlp2

Similar

Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource LanguagesMultilingual Multimodal Pre-Training with TLI for Zero-Shot Image-Text Retrieval in Low-Resource African LanguagesAdaptive Scheduling for Multi-Task LearningIntermediate-Task Training Effects on Zero-Shot XTREME Classification Across LanguagesActive Learning with Multiple Annotations for Comparable Data Classification TaskMultimodal vs. Text-Only Pre-Training for Cross-Lingual NER in Low-Resource Languages

Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages

Training deep learning networks with minimal supervision has gained significant research attention d

Multilingual Multimodal Pre-Training with TLI for Zero-Shot Image-Text Retrieval in Low-Resource African Languages

This paper studies zero-shot cross-lingual transfer of vision-language models. Specifically, we focu

Adaptive Scheduling for Multi-Task Learning

To train neural machine translation models simultaneously on multiple tasks (languages), it is commo

Intermediate-Task Training Effects on Zero-Shot XTREME Classification Across Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Active Learning with Multiple Annotations for Comparable Data Classification Task

Supervised learning algorithms for identifying comparable sentence pairs from a dominantly non-paral

Multimodal vs. Text-Only Pre-Training for Cross-Lingual NER in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident