Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

GIZ AI4D Africa Language Challenge - Round 2

Domaine:

natural language processing

Type de record:

dataset

Calling on the Zindi community to help uncover and create African Language Datasets for improved representation in the field of NLP
There is no data for this competition.
You can download:
AI4D_Documentation.docx - This contains guidelines for your documentation. It includes questions on motivation, composition, collection process, recommended uses, and so on.
When you make a submission you will need to submit a data set in the same format as AI4D_Data.txt, along with the dataset's documentation. The documentation should answer the questions in AI4D_Documentation.docx. This challenge calls on you to submit African language datasets (annotated or otherwise) that are representative and balanced and useful for downstream NLP tasks.For your submission to be eligible, the data must meet the following criteria:
The languages must be indigenous to Uganda, Ghana, or South Africa
Data should be sentence split and not tokenized
Each dataset submission must be accompanied by a dataset documentation.
The documentation covers the motivation, composition, collection process, recommended uses, and so on. See this paper for further details. A report template is provided.
While you can adapt the sections covered in the documentation, you must include the final section- an explanation on how you would expand this data set if you won the $1,500 research grant.
Our intention is that the datasets are kept free and open for public use under a Creative Commons 4.0 license or similar. Data already licensed under more restrictive terms will not be eligible
You must upload your submission to the competition one file at a time. You must include documentation for each submission that describes the submitted dataset. Note that there will be no scores on this leaderboard. If you make multiple submissions, each of your submissions will be judged independently of each other. Up to three submissions from an individual or team will be considered. It is possible for someone to win multiple prizes.
You should provide two files for each submission:
ONE txt file with the language data (or multiple files in the case of multilingual datasets)
ONE pdf file accompanying the datasheet that documents its motivation, composition, collection process, recommended uses, and so on. See this paper for further details.
Please label your files:
username_data_XXX.txt
username_documentation_XXX.pdf
Where XXX is a unique ID to indicate which datasheet goes with which documentation if you make multiple submissions. Note that you can also zip the files.

Visit

zindi.africa

Tags

competitionzindicollectionresearch

Similaires

AI4D -- African Language Dataset Challengetheyorubayesian/AI4D-Dataset-ChallengeAI4D Malawi News Classification ChallengeAI4D Yorùbá Machine Translation ChallengeEKivutha/-AI4D-Tourism-Classification-ChallengeAI4D Takwimu Lab - Machine Translation Challenge

AI4D -- African Language Dataset Challenge

As language and speech technologies become more advanced, the lack of fundamental digital resources

theyorubayesian/AI4D-Dataset-Challenge

AI4D has put out an African Languages Dataset Challenge & I will be taking the opportunity to brush

AI4D Malawi News Classification Challenge

Can you classify Malawi news articles in Chichewa?
The data was collected from news publications in Malawi. tNyasa Ltd Data Science Lab have used three main broadcasters: the Nation Online newspaper, Radio Maria and the Malawi Broadcasting Corporation. The

AI4D Yorùbá Machine Translation Challenge

Can you translate Yorùbá to English?
The training data consist of 10,054 parallel Yorùbá-English sentences from different domains like news, Yorùbá proverbs, movie transcript, ted talks, radio broadcast transcript, localization translation, and books.
Va

EKivutha/-AI4D-Tourism-Classification-Challenge

Use tourism survey data and ML to classify the range of expenditures a tourist spends in Tanzania

AI4D Takwimu Lab - Machine Translation Challenge

Can you translate French to Fongbe and Ewe?
This is a parallel corpus dataset for machine translation from French to Ewe and French to Fongbe, languages from Togo and Benin respectively. It contains roughly 23 000 French to Ewe and 53 000 French to Fongbe p