Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

GIZ Kinyarwanda Text Cleaning and Augmentation Competition

Domain:

natural language processing

Record type:

dataset

Can you collect and curate English-Kinyarwanda text data for machine translation tasks?
Approximately 55 parallel Kinyarwanda-English sentences will be provided for data cleaning along. This data is to help you get started on this process.
You are tasked with finding additional data sources of parallel Kinyarwanda-English sentences. You need to clearly document where and how you downloaded the data, however it is preferable that you input the data straight into your script using an API.
GIZ is particularly interested in domain-specific text data from fields such as health, agriculture, tourism, etc,. We recommend that you focus on one field, you can even specialize in a specific subfield if you’d like. As quantity and quality of data will result in a strong model.
Please do not use the JW300 datasets as this may skew the distribution of data across fields.
The objective of this challenge is to create a script that will clean Kinyarwanda-English parallel sentences.
You are encouraged to use a rules-based approach along with machine learning if you think it is applicable. Remember to consider your script’s efficiency and memory usage during execution.
You are welcome to create a machine translation but it will not add to your final score.

Visit

zindi.africa

Tasks

machine translation

Languages

Kinyarwanda

Tags

competitionzindicollectionresearch

Similar

Kinyarwanda-speech-to-text-ASR/Kinyarwanda-speech-to-text-ASRkinyarwanda text corpusLLM-Driven Text Augmentation across Media and LanguagesReal-time recognition and translation of Kinyarwanda sign language into Kinyarwanda textOptical Character Recognition and text cleaning in the indigenous South African languagesLanguage-Independent Data Augmentation for Text Classification [LiDA]

Kinyarwanda-speech-to-text-ASR/Kinyarwanda-speech-to-text-ASR

Kinyarwanda Automatic Speech Recognition Model based on Whisper # KinyaWhisper - Kinyarwanda Automa

kinyarwanda text corpus

LLM-Driven Text Augmentation across Media and Languages

The proliferation of fake news across social media, headlines, and news articles poses major challen

Real-time recognition and translation of Kinyarwanda sign language into Kinyarwanda text

Optical Character Recognition and text cleaning in the indigenous South African languages

This article represents follow-up work on unpublished presentations by the authors of text and corpu

Language-Independent Data Augmentation for Text Classification [LiDA]

Building high-performance text classification models in low-resource languages is a challenging task