Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
ShaElsVas
Host:avatar
Most social media users come from the Global South, where harmful content usually appears in local languages. Yet, AI-driven moderation systems struggle with low-resource languages spoken in these regions. Through semi-structured interviews with 22 AI experts working on harmful content detection in four low-resource languages: Tamil (South Asia), Swahili (East Africa), Maghrebi Arabic (North Africa), and Quechua (South America)--we examine systemic issues in building automated moderation tools for these languages. Our findings reveal that beyond data scarcity, socio-political factors such as tech companies' monopoly on user data and lack of investment in moderation for low-profit Global South markets exacerbate historic inequities. Even if more data were available, the English-centric and data-intensive design of language models and preprocessing techniques overlooks the need to design for morphologically complex, linguistically diverse, and code-mixed languages. We argue these limitations are not just technical gaps caused by "data scarcity" but reflect structural inequities, rooted in colonial suppression of non-Western languages. We discuss multi-stakeholder approaches to strengthen local research capacity, democratize data access, and support language-aware solutions to improve automated moderation for low-resource languages. Accepted to AIES 2025

Visit

arxiv.org

Tasks

hate speech detectiontext classification

Languages

Arabic, Libyan SpokenArabic, Moroccan SpokenSwahili

Tags

Computation and LanguageHuman-Computer Interaction

Similar

The Paradox of Undetected Harm: Content Moderation Blind Spots in Low-Resource LanguagesContent Moderation in the Global South: A Comparative Study of Four Low-Resource LanguagesOptimizing Data Pipelines for Low-Resource Settings: A Case Study from Community Health in KenyaAuditing YouTube Content Moderation in Low Resource Language SettingsThe eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource LanguagesThe Usefulness of Imperfect Speech Data for ASR Development in Low-Resource Languages

The Paradox of Undetected Harm: Content Moderation Blind Spots in Low-Resource Languages

This paper explores a systemic paradox in global content moderation: harmful content in low-resource

Content Moderation in the Global South: A Comparative Study of Four Low-Resource Languages

Over the past 18 months, the Center for Democracy and Technology (CDT) has been studying how content

Optimizing Data Pipelines for Low-Resource Settings: A Case Study from Community Health in Kenya

ABSTRACT

Auditing YouTube Content Moderation in Low Resource Language Settings

While there has been increasing attention paid to the potential harms perpetuated by online platform

The eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource Languages

Efficiently and accurately translating a corpus into a low-resource language remains a challenge, re

The Usefulness of Imperfect Speech Data for ASR Development in Low-Resource Languages

When the National Centre for Human Language Technology (NCHLT) Speech corpus was released, it create