Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings

Domain:

natural language processing

Record type:

paperdataset
Creator:
GaiSonLeeKo,
Host:avatar
Content moderation research has recently made significant advances, but remains limited in serving the majority of the world's languages due to the lack of resources, leaving millions of vulnerable users to online hostility. This work presents a large-scale human-annotated multi-task benchmark dataset for abusive language detection in Tigrinya social media with joint annotations for three tasks: abusiveness, sentiment, and topic classification. The dataset comprises 13,717 YouTube comments annotated by nine native speakers, collected from 7,373 videos with a total of over 1.2 billion views across 51 channels. We developed an iterative term clustering approach for effective data selection. Recognizing that around 64% of Tigrinya social media content uses Romanized transliterations rather than native Ge'ez script, our dataset accommodates both writing systems to reflect actual language use. We establish strong baselines across the tasks in the benchmark, while leaving significant challenges for future contributions. Our experiments demonstrate that small fine-tuned models outperform prompted frontier large language models (LLMs) in the low-resource setting, achieving 86.67% F1 in abusiveness detection (7+ points over best LLM), and maintain stronger performance in all other tasks. The benchmark is made public to promote research on online safety. Accepted at NeurIPS 2025

Visit

arxiv.org

Tasks

hate speech detectionsentiment analysistext classificationtopic classification

Languages

Tigrigna

Tags

Computation and LanguageI.2.7

Similar

Multi-Hall-SA: A Cross-lingual Benchmark for Multi-Type Hallucination Detection in Low-Resource South African LanguagesMulti-source Intermediate-task Training for Low-resource XTREME Language GeneralizationEl-amin/FairCXRnet-A-Multi-Task-Learning-Model-for-Chest-X-Ray-Classification-for-Low-Resource-SettingsA multi-task learning framework for sentiment analysis and news classification for low-resource languageEnglish Intermediate-Task Training for Robustness in Low-Resource XTREME Benchmark LanguagesAbusive Content Detection in Arabic Tweets Using Multi-Task Learning and Transformer-Based Models

Multi-Hall-SA: A Cross-lingual Benchmark for Multi-Type Hallucination Detection in Low-Resource South African Languages

Multi-source Intermediate-task Training for Low-resource XTREME Language Generalization

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

El-amin/FairCXRnet-A-Multi-Task-Learning-Model-for-Chest-X-Ray-Classification-for-Low-Resource-Settings

This is notebook isfrom a research paper "FairCXRnet: A Multi-Task Learning Model for Domain Adaptat

A multi-task learning framework for sentiment analysis and news classification for low-resource language

Despite the growing progress in Natural Language Processing (NLP), low-resource languages such as Ha

English Intermediate-Task Training for Robustness in Low-Resource XTREME Benchmark Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Abusive Content Detection in Arabic Tweets Using Multi-Task Learning and Transformer-Based Models

Different social media platforms have become increasingly popular in the Arab world in recent years.