Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Kencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
WanWanzare, LilianIndMcO
Hôte:avatar
Indigenous African languages are categorized as under-served in Natural Language Processing. They therefore experience poor digital inclusivity and information access. The processing challenge with such languages has been how to use machine learning and deep learning models without the requisite data. The Kencorpus project intends to bridge this gap by collecting and storing text and speech data that is good enough for data-driven solutions in applications such as machine translation, question answering and transcription in multilingual communities. The Kencorpus dataset is a text and speech corpus for three languages predominantly spoken in Kenya: Swahili, Dholuo and Luhya. Data collection was done by researchers from communities, schools, media, and publishers. The Kencorpus' dataset has a collection of 5,594 items - 4,442 texts (5.6M words) and 1,152 speech files (177hrs). Based on this data, Part of Speech tagging sets for Dholuo and Luhya (50,000 and 93,000 words respectively) were developed. We developed 7,537 Question-Answer pairs for Swahili and created a text translation set of 13,400 sentences from Dholuo and Luhya into Swahili. The datasets are useful for downstream machine learning tasks such as model training and translation. We also developed two proof of concept systems: for Kiswahili speech-to-text and machine learning system for Question Answering task, with results of 18.87% word error rate and 80% Exact Match (EM) respectively. These initial results give great promise to the usability of Kencorpus to the machine learning community. Kencorpus is one of few public domain corpora for these three low resource languages and forms a basis of learning and sharing experiences for similar works especially for low resource languages. 24 pages, 6 figures

Visit

arxiv.org

Tasks

automatic speech recognitionmachine translationspeech processing

Languages

DholuoLuhyaSwahiliSwahili, CoastalSwahili, Congo

Tags

Computation and Language

Similaires

Kencorpus: Kenyan Languages Corpus for Machine Learning and Natural Language ProcessingKencorpus: Kenyan Languages Corpus for Natural Language Processing and Machine LearningReplication Data for Igbo Natural Language Processing Tasks I Igbo Synchronised Corpus for Natural Language Processing TasksReplication Data for Igbo Natural Language Processing Tasks II Igbo Synchronised Corpus for Natural Language Processing TasksVoicing the Voiceless: Developing a Dholuo Parallel Corpus for Natural Language ProcessingReplication Data for Igbo Natural Language Processing Tasks

Kencorpus: Kenyan Languages Corpus for Machine Learning and Natural Language Processing

This project collected text and speech corpora for three languages in Kenya: Kiswahili, Dholuo and 3 Luhya dialects (Lumarachi, Logooli and Lubukusu). Primary data was collected from the respective language communities, which included Indigenous stories and narrati

Kencorpus: Kenyan Languages Corpus for Natural Language Processing and Machine Learning

Kencorpus: Kenyan Languages Corpus for Natural Language Processing and Machine Learning

Poster presented at the Deep Learning Indaba 2022 by Lillian Wanzare

Replication Data for Igbo Natural Language Processing Tasks I Igbo Synchronised Corpus for Natural Language Processing Tasks

The Igbo synchronised corpus (IgboSynCorp) is an annotated corpus of spoken Igbo created by a team o

Replication Data for Igbo Natural Language Processing Tasks II Igbo Synchronised Corpus for Natural Language Processing Tasks

The Igbo synchronised corpus (IgboSynCorp) is an annotated corpus of spoken Igbo created by a team o

Voicing the Voiceless: Developing a Dholuo Parallel Corpus for Natural Language Processing

Despite the exponential growth of Natural Language Processing (NLP) technologies worldwide, African

Replication Data for Igbo Natural Language Processing Tasks

The Igbo synchronised corpus (IgboSynCorp) is an annotated corpus of spoken Igbo created by a team of linguists and NLP experts at the University of Ibadan and Afe Babalola University, Nigeria. The project was designed to create an open access labelled and unlabell