Logo Lanfrica

Copyright and African NLP: What we can learn from the practices and activities of Africa-focused natural language processing

Domain:

natural language processing

Record type:

paper
Creator:
Okorie, Chijioke
Publisher:
Zenodo
Host:avatar
This paper reports pilot empirical work examining a sample of African NLP research papers published between 2019 and 2024 along these dimensions: the type of resource created, the data sources used to create them, the licences attached, the funders and supporters behind the work, and the repositories on which they appear. It presents preliminary findings, notably that the data sources behind African NLP are predominantly copyright-protected materials from news media and personal information; that the resources produced are themselves frequently released without clear or explicit licences; and that an overwhelming number of hosting and funding and investments for this work is from outside the continent.