Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kenya Legal NLP Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
gma
Host:
Named entity recognition annotations over Kenya legal documents in English and Swahili. Covers the Constitution of Kenya 2010, Employment Act, Land Act, and Data Protection Act. License: CC BY 4.0Author: Gabriel Mahia — gabrielmahia.github.ioSource: Public domain Kenya government legal documents Research Context

Visit

huggingface.co

Tasks

information extractionnamed entity recognition

Languages

Swahili

Tags

kenyalegalnlpnerconstitutionemployment-lawland-laweast-africa

Licenses

cc-by-4.0

Similar

isiXhosa NLP DatasetAmazigh/Berber NLP DatasetSwahili Civic NLP DatasetSilva3012/xhosa-nlp-datasetMahammadRiyazShek/multilingual-nlp-dataset-processorAmharic Multi-Task NLP Dataset

isiXhosa NLP Dataset

A high-quality, comprehensive isiXhosa (Xhosa) NLP training dataset carefully collected, cleaned, an

Amazigh/Berber NLP Dataset

Ghosts in the Font: Reviving a Low-Resource Language for AI A structured NLP dataset and toolkit for

Swahili Civic NLP Dataset

Labelled Swahili text samples for civic, health, financial, and educational NLP tasks in East Africa

Silva3012/xhosa-nlp-dataset

My attempt at building a xhosa NLP dataset --- language: - xh - en license: other multilinguality:

MahammadRiyazShek/multilingual-nlp-dataset-processor

End-to-end NLP pipeline processing multilingual datasets across 38 languages and 6+ locales — 10,500

Amharic Multi-Task NLP Dataset