Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages

Domain:

natural language processing

Record type:

paper
Creator:
van
Publisher:
arXiv
Host:avatar
Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are rarely applied. CC-BY-SA and CC-BY-NC cannot be combined in a single published dataset; a NoDerivs clause silently prohibits tokenisation and annotation. This paper audits the license provenance of over twenty corpus families used in African NLP, constructs a six-tier compatibility matrix, and applies it to three case-study languages: Kituba/Munukutuba, Zarma, and Moore. Four failure modes are documented with primary-source evidence: outright prohibition (JW300, removed from OPUS after a legal audit confirmed Terms of Service violation); composite license misrepresentation (WAXAL, whose CC-BY 4.0 claim is contradicted by its own HuggingFace dataset card); a NoDerivs clause hidden behind a CC-BY label (Tanzil); and data persistence failure (the Congolese Radio Corpus, where 402 of 405 source URLs are now dead). A pre-annotation due diligence checklist and a survey of legally clean enrichment opportunities close the paper. 12 pages. Published in Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC-COLING 2026, pages 128-139

Visit

doi.orgarxiv.org

Languages

KitubaKitubaZarma

Tags

Computation and Language (cs.CL)FOS: Computer and information sciencesI.2.7

Licenses

Creative Commons Attribution Non Commercial Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-sa/4.0/legalcode

Similar

Building Corpora for Low-Resource Kenyan LanguagesGoogle Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An OverviewTransformers for Low-Resource African Languages Transformer Architectures and the Self-Attention Mechanism for Low-Resource African Languages: A Survey of Approaches, Benchmarks, and Open ChallengesChatGPT MT: Competitive for High- (but not Low-) Resource LanguagesGhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian LanguagesDynAg Open Voice Dataset for Low-Resource Bihari Languages

Building Corpora for Low-Resource Kenyan Languages

Natural Language Processing is a crucial frontier in artificial intelligence, with broad application

Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An Overview

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and automatic speech recognition

Transformers for Low-Resource African Languages Transformer Architectures and the Self-Attention Mechanism for Low-Resource African Languages: A Survey of Approaches, Benchmarks, and Open Challenges

ChatGPT MT: Competitive for High- (but not Low-) Resource Languages

Large language models (LLMs) implicitly learn to perform a range of language tasks, including machin

GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages

Low resource languages present unique challenges for natural language processing due to the limited

DynAg Open Voice Dataset for Low-Resource Bihari Languages

jjkhkjhkjh