Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Extreme Multi-Label Text Classification for Less-Represented Languages and Low-Resource Environments: Advances and Lessons Learned

Domaine:

natural language processing

Type de record:

paper
Créateur:
NikBlaBosSen
Éditeur:
MDP
Hôte:
Amid ongoing efforts to develop extremely large, multimodal models, there is increasing interest in efficient Small Language Models (SLMs) that can operate without reliance on large data-centre infrastructure. However, recent SLMs (e.g., LLaMA or Phi) with up to three billion parameters are predominantly trained in high-resource languages, such as English, which limits their applicability to industries that require robust NLP solutions for less-represented languages and low-resource settings, particularly those requiring low latency and adaptability to evolving label spaces. This paper examines a retrieval-based approach to multi-label text classification (MLC) for a media monitoring dataset, with a particular focus on less-represented languages, such as Slovene. This dataset presents an extreme MLC challenge, with instances labelled using up to twelve thousand categories. The proposed method, which combines retrieval with computationally efficient prediction, effectively addresses challenges related to multilinguality, resource constraints, and frequent label changes. We adopt a model-agnostic approach that does not rely on a specific model architecture or language selection. Our results demonstrate that techniques from the extreme multi-label text classification (XMC) domain outperform traditional Transformer-based encoder models, particularly in handling dynamic label spaces without requiring continuous fine-tuning. Additionally, we highlight the effectiveness of this approach in scenarios involving rare labels, where baseline models struggle with generalisation.

Visit

doi.org

Tasks

text classification

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Advances in Low-Resource and Endangered LanguagesAssessing BERT-based models for Arabic and low-resource languages in crime text classificationFlick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource LanguagesText Normalization for Low Resource LanguagesText Classification Based on Convolutional Neural Networks and Word Embedding for Low-Resource Languages: TigrinyaHaEmoC-MLTC: Hausa Emotion Corpus for Multi-Label Text Classification from Twitter

Advances in Low-Resource and Endangered Languages

This paper reports on the approaches and results for the collection, analysis, and processing of low

Assessing BERT-based models for Arabic and low-resource languages in crime text classification

The bidirectional encoder representations from Transformers (BERT) has recently attracted considerab

Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages

Training deep learning networks with minimal supervision has gained significant research attention d

Text Normalization for Low Resource Languages

This repository contains code related to the Google open source internship project Text Normalization for Low Resource Languages.

Text Classification Based on Convolutional Neural Networks and Word Embedding for Low-Resource Languages: Tigrinya

This article studies convolutional neural networks for Tigrinya (also referred to as Tigrigna), whic

HaEmoC-MLTC: Hausa Emotion Corpus for Multi-Label Text Classification from Twitter

The dataset comprises 12,761 Hausa tweets, each annotated with 11 distinct emotions: anger, sadness,