Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kencorpus: Kenyan Languages Corpus for Machine Learning and Natural Language Processing

Domain:

natural language processing

Record type:

dataset
This project collected text and speech corpora for three languages in Kenya: Kiswahili, Dholuo and 3 Luhya dialects (Lumarachi, Logooli and Lubukusu). Primary data was collected from the respective language communities, which included Indigenous stories and narratives from student compositions, native language media stations, and publishers – in order to include genres of texts representative of everyday language use in the communities. A total of 4,442 texts were collected: 2909 for Swahili, 546 texts for Dholuo, 483 texts for Lumarachi, 135 texts for Lubukusu, and 359 texts for Logooli. A total of 1,152 files containing spontaneous speech data were collected, which total to 176 hours, 29 minutes, and 46 seconds: 104 files (19 hours, 10 minutes, 57 seconds) for Swahili, 512 files (99 hours, 3 minutes, 8 seconds) for Dholuo, 138 files (15 hours, 37 minutes, 46 seconds) for Lumarachi, 354 files (30 hours, 11 minutes) for Lubukusu, and annotated 44 files (12 hours, 26 minutes, 55 seconds) for Lulogooli.

Visit

dataverse.harvard.edu

Tasks

automatic speech recognitionmachine translationtext to speechspeech translation

Languages

BukusuDholuoLulogooliSwahiliSwahili, CoastalSwahili, Congo

Tags

lacuna fundfair forward

Similar

Kencorpus: Kenyan Languages Corpus for Natural Language Processing and Machine LearningKencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing TasksKencorpus: Kenyan Languages Corpus

Kencorpus: Kenyan Languages Corpus for Natural Language Processing and Machine Learning

Kencorpus: Kenyan Languages Corpus for Natural Language Processing and Machine Learning

Poster presented at the Deep Learning Indaba 2022 by Lillian Wanzare

Kencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks

Indigenous African languages are categorized as under-served in Natural Language Processing. They th

Kencorpus: Kenyan Languages Corpus

This project collected text and speech corpora for Languages in Kenya. In KenCorpus project, three l