Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
van
Éditeur:
arXiv
Hôte:avatar
Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are rarely applied. CC-BY-SA and CC-BY-NC cannot be combined in a single published dataset; a NoDerivs clause silently prohibits tokenisation and annotation. This paper audits the license provenance of over twenty corpus families used in African NLP, constructs a six-tier compatibility matrix, and applies it to three case-study languages: Kituba/Munukutuba, Zarma, and Moore. Four failure modes are documented with primary-source evidence: outright prohibition (JW300, removed from OPUS after a legal audit confirmed Terms of Service violation); composite license misrepresentation (WAXAL, whose CC-BY 4.0 claim is contradicted by its own HuggingFace dataset card); a NoDerivs clause hidden behind a CC-BY label (Tanzil); and data persistence failure (the Congolese Radio Corpus, where 402 of 405 source URLs are now dead). A pre-annotation due diligence checklist and a survey of legally clean enrichment opportunities close the paper. 12 pages. Published in Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC-COLING 2026, pages 128-139

Visit

doi.orgarxiv.org

Languages

KitubaKitubaZarma

Tags

Computation and Language (cs.CL)FOS: Computer and information sciencesI.2.7

Licenses

Creative Commons Attribution Non Commercial Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-sa/4.0/legalcode

Similaires

Building Corpora for Low-Resource Kenyan LanguagesGoogle Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An OverviewTransformers for Low-Resource African Languages Transformer Architectures and the Self-Attention Mechanism for Low-Resource African Languages: A Survey of Approaches, Benchmarks, and Open ChallengesChatGPT MT: Competitive for High- (but not Low-) Resource LanguagesGhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian LanguagesDynAg Open Voice Dataset for Low-Resource Bihari Languages

Building Corpora for Low-Resource Kenyan Languages

Natural Language Processing is a crucial frontier in artificial intelligence, with broad application

Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An Overview

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and automatic speech recognition

Transformers for Low-Resource African Languages Transformer Architectures and the Self-Attention Mechanism for Low-Resource African Languages: A Survey of Approaches, Benchmarks, and Open Challenges

ChatGPT MT: Competitive for High- (but not Low-) Resource Languages

Large language models (LLMs) implicitly learn to perform a range of language tasks, including machin

GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages

Low resource languages present unique challenges for natural language processing due to the limited

DynAg Open Voice Dataset for Low-Resource Bihari Languages

jjkhkjhkjh