Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Building Text and Speech Benchmark Datasets and Models for Low‐Resourced East African Languages: Experiences and Lessons

Domain:

natural language processing

Record type:

paperdatasetmodel
Creator:
Joyce, Nakatumba-NabendeClaire, BabiryePetJer
Publisher:
WILEY
Host:
ABSTRACT Africa has over 2000 languages; however, those languages are not well represented in the existing natural language processing ecosystem. African languages lack essential digital resources to effectively engage in advancing language technologies. There is a need to generate high‐quality natural language processing resources for low‐resourced African languages. Obtaining high‐quality speech and text data is expensive and tedious because it can involve manual sourcing and verification of data sources. This paper discusses the process taken to curate and annotate text and speech datasets for five East African languages: Luganda, Runyankore‐Rukiga, Acholi, Lumasaba, and Swahili. We also present results obtained from baseline models for machine translation, topic modeling and classification, sentiment classification, and automatic speech recognition tasks. Finally, we discuss the experiences, challenges, and lessons learned in creating the text and speech datasets.

Visit

doi.org

Tasks

automatic speech recognitionmachine translationsentiment analysisspeech processingtext classification

Languages

AcholiChigaGandaNyankoreSwahili

Licenses

http://creativecommons.org/licenses/by/4.0/

Similar

Building Text and Speech Datasets for Low Resourced Languages: A Case of Languages in East AfricaBuilding Text-to-Speech Models for Low-Resourced Languages from Crowdsourced DataQuestion-Answering in a Low-resourced Language: Benchmark Dataset and Models for TigrinyaDatasets Collection Framework for Low-Resourced Languages in South AfricaEnd-to-End Text-To-Speech synthesis for under resourced South African languagesCombining Unsupervised and Text Augmented Semi-Supervised Learning for Low Resourced Autoregressive Speech Recognition

Building Text and Speech Datasets for Low Resourced Languages: A Case of Languages in East Africa

Africa has over 2000 languages; however, those languages are not well represented in the existing Natural Language Processing ecosystem. African languages lack essential digital resources to be engaged effectively in the advancing language technologies. This growin

Building Text-to-Speech Models for Low-Resourced Languages from Crowdsourced Data

Text-to-speech (TTS) models have expanded the scope of digital inclusivity by becoming a basis for a

Question-Answering in a Low-resourced Language: Benchmark Dataset and Models for Tigrinya

Question-Answering (QA) has seen significant advances recently, achieving near human-level performan

Datasets Collection Framework for Low-Resourced Languages in South Africa

End-to-End Text-To-Speech synthesis for under resourced South African languages

Combining Unsupervised and Text Augmented Semi-Supervised Learning for Low Resourced Autoregressive Speech Recognition

Recent advances in unsupervised representation learning have demonstrated the impact of pretraining