Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Content Moderation in the Global South: A Comparative Study of Four Low-Resource Languages

Domain:

natural language processingdigital infrastructure

Record type:

paper
Creator:
MonAliDha
Publisher:
Cen
Host:
Over the past 18 months, the Center for Democracy and Technology (CDT) has been studying how content moderation systems operate across multiple regions in the Global South, with a focus on South Asia, North and East Africa, and South America. Our team studied four languages: the different Maghrebi Arabic Dialects (Elswah, 2024a), Kiswahili (Elswah, 2024b), Tamil (Bhatia & Elswah, 2025), and Quechua (Thakur, 2025). These languages and dialects are considered “low resource” due to the scarcity of training data available to develop equitable and accurate AI models for them. We did this through essential collaborations with regional civil society organizations in the Global South to help us understand the local dynamics of their digital environments. Content moderation remains an area that technology companies keep largely inaccessible to public scrutiny, except for the information they choose to disclose. Our findings significantly contribute to the scientific and policy communities’ understanding of content moderation and its challenges in the Global South. The data we present in this report also contributes to our understanding of the information environment in the Global South, which is understudied in current scholarship.

Visit

doi.org

Languages

Arabic, Libyan SpokenArabic, Moroccan SpokenSwahiliSwahili, CoastalSwahili, Congo

Licenses

https://creativecommons.org/licenses/by/4.0/legalcode

Similar

The Paradox of Undetected Harm: Content Moderation Blind Spots in Low-Resource LanguagesAuditing YouTube Content Moderation in Low Resource Language SettingsThink Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource LanguagesInequalities and content moderationBLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource LanguagesA comparative study of the locative system in South-Tanzanian Bantu languages

The Paradox of Undetected Harm: Content Moderation Blind Spots in Low-Resource Languages

This paper explores a systemic paradox in global content moderation: harmful content in low-resource

Auditing YouTube Content Moderation in Low Resource Language Settings

While there has been increasing attention paid to the potential harms perpetuated by online platform

Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages

Most social media users come from the Global South, where harmful content usually appears in local l

Inequalities and content moderation

As the harms of hate speech, mis/disinformation and incitement to violence on social media have beco

BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages

Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain co

A comparative study of the locative system in South-Tanzanian Bantu languages

The paper presents a comparative analysis of locative expressions in four South-Tanzanian Bantu lang