Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

An Open Source System for Crowd Sourcing an African Language Short Story Corpus

Domain:

natural language processing

Record type:

paper
Many African languages are under resourced in having open access corpora for use in developing technological applications such as grammar checkers, spell checkers, speech to text, text to speech and machine translation tools. This may lead to a decline in all cultural traits associated with the peoples that speak these languages. To enable collection of textual corpora and long term preservation of positive cultural characteristics, the design considerations and implementation of an open source online short story competition collection and evaluation system are described. The system is written in PHP and can be relatively cheaply deployed on shared hosting servers available from many African hosting providers. This allows for the possibility of a decentralized collection of stories, as well as adaptation and improvements of the software to different types of short story competitions. The software has been used for two short story competitions across the African continent with the aim of providing stories suitable for children. Holding the competition online has enabled participation from a wide variety of locations, but most of the submissions have came from African countries with relatively good information technology infrastructure. Preparation for a third competition is in progress.

Visit

doi.orgtuvutepamoja.africatuvutepamoja.africa

Tags

railTuvute Pamoja

Similar

Developing an Open-Source Corpus of Yoruba SpeechAlkhalil Corpus: An Open-Source Thematic and Lemmatized Corpus for Modern Standard ArabicSagalee: An Open Source ASR Dataset for Oromo LanguageSagalee: An Open Source ASR Dataset for Oromo LanguageAn Open-Source Monitoring System for Remote Solar Power ApplicationsAn Open Source Biometric Patient Identification System for Low Resource Setting

Developing an Open-Source Corpus of Yoruba Speech

This paper introduces an open-source speech dataset for Yoruba - one of the largest low-resource West African languages spoken by at least 22 million people. Yoruba is one of the official languages of Nigeria, Benin and Togo, and is spoken in other neighboring Afri

Alkhalil Corpus: An Open-Source Thematic and Lemmatized Corpus for Modern Standard Arabic

The availability of large annotated corpora remains a major challenge for the development of natural

Sagalee: An Open Source ASR Dataset for Oromo Language

Sagalee is Speech Recognition Dataset for Oromo language Presented in the paper: Sagalee: an Open So

Sagalee: An Open Source ASR Dataset for Oromo Language

Sagalee is Speech Recognition Dataset for Oromo language Presented in the paper: Sagalee: an Open So

An Open-Source Monitoring System for Remote Solar Power Applications

Renewable energy systems are an increasingly popular way to generate electricity. As with any new te

An Open Source Biometric Patient Identification System for Low Resource Setting

It is estimated that as many as 1.5 billion people globally do not possess any form of identificatio