Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
Mbogho, AudreyAwuor, QuinKipkebut, AndrewWanzare, Lilian
Hôte:avatar
Natural Language Processing is a crucial frontier in artificial intelligence, with broad applications in many areas, including public health, agriculture, education, and commerce. However, due to the lack of substantial linguistic resources, many African languages remain underrepresented in this digital transformation. This paper presents a case study on the development of linguistic corpora for three under-resourced Kenyan languages, Kidaw'ida, Kalenjin, and Dholuo, with the aim of advancing natural language processing and linguistic research in African communities. Our project, which lasted one year, employed a selective crowd-sourcing methodology to collect text and speech data from native speakers of these languages. Data collection involved (1) recording conversations and translation of the resulting text into Kiswahili, thereby creating parallel corpora, and (2) reading and recording written texts to generate speech corpora. We made these resources freely accessible via open-research platforms, namely Zenodo for the parallel text corpora and Mozilla Common Voice for the speech datasets, thus facilitating ongoing contributions and access for developers to train models and develop Natural Language Processing applications. The project demonstrates how grassroots efforts in corpus building can support the inclusion of African languages in artificial intelligence innovations. In addition to filling resource gaps, these corpora are vital in promoting linguistic diversity and empowering local communities by enabling Natural Language Processing applications tailored to their needs. As African countries like Kenya increasingly embrace digital transformation, developing indigenous language resources becomes essential for inclusive growth. We encourage continued collaboration from native speakers and developers to expand and utilize these corpora. 13 pages, 1 figure, intend to submit to a Springer Nature journal

Visit

arxiv.org

Languages

DawidaDholuoKalenjinKipsigisSwahiliSwahili, CoastalSwahili, Congo

Tags

Computation and Language

Similaires

waleghwa/low-resource-language-data: Parallel Corpora for Kiswahili and Kidaw'ida, Kalenjin and DholuoBuilding Corpora for Low-Resource Kenyan LanguagesEnhancing Pos Tagging For Low-Resource Languages: A Case Study On DholuoBuilding and Evaluating Somali Language CorporaOpen Kalenjin Automatic Speech Recognition: Adapting Parakeet-TDT to a Low-Resource Nilotic LanguageHeeLeeOss/low-resource-parallel-corpora

waleghwa/low-resource-language-data: Parallel Corpora for Kiswahili and Kidaw'ida, Kalenjin and Dholuo

Description: The dataset consists of three parallel corpora: K

Building Corpora for Low-Resource Kenyan Languages

Natural Language Processing is a crucial frontier in artificial intelligence, with broad application

Enhancing Pos Tagging For Low-Resource Languages: A Case Study On Dholuo

Building and Evaluating Somali Language Corpora

Open Kalenjin Automatic Speech Recognition: Adapting Parakeet-TDT to a Low-Resource Nilotic Language

Kalenjin, the Highland Nilotic language cluster of Kenya's third-largest ethnic community (6.36 mill

HeeLeeOss/low-resource-parallel-corpora

Openly-licensed, consented parallel corpora for under-served languages — with a clear schema, consen