Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

KenPOS

Domain:

natural language processing

Record type:

dataset
Creator:
And
Host:
KenPOS is a part-of-speech (POS) tagged corpus for Kenyan languages, featuring 156,994 tokens across four languages. The dataset provides manually annotated POS tags for low-resource Kenyan languages, enabling NLP research and applications. Language Code Tokens Sentences Files Unique POS Tags Dholuo dho 54,712 70 168 114 Lubukusu lbk 51,900 154 62

Visit

huggingface.co

Tasks

part of speech tagging

Languages

BukusuDholuoLulogooliOlumarachi

Tags

kenyan-languagesdholuolubukusulumarachilulogoolipos-tagginglow-resource-languagesafrican-languages

Licenses

cc-by-4.0

Similar

KenPOSKenPos: Kenyan Languages Part of Speech Tagged dataset

KenPOS

KenPOS is a part-of-speech (POS) tagged corpus for Kenyan languages, featuring 156,994 tokens across

KenPos: Kenyan Languages Part of Speech Tagged dataset

This project developed a Part of Speech (POS) Tagged dataset of 2 languages in Kenya: Dholuo and 3 Luhya dialects (Lumarachi, Lulogooli, and Lubukusi). The project tagged approximately 143,000 words, which includes about 50,000 words for Dholuo, 27,900 words for Lu