Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TaTA: A Multilingual Table-to-Text Dataset for African Languages

Domaine:

natural language processinghealthcare

Type de record:

dataset

TaTA (Table-to-Text in African languages) is the first large multilingual table-to-text datasets with a focus on African languages. The dataset is parallel and covers nine languages, eight of which are spoken in Africa: Arabic, English, French, Hausa, Igbo, Portuguese, Swahili, Yorùbá, and Russian.

TaTa was created by extracting tables from charts in 71 PDF reports published between 1990 and 2021 by the Demographic and Health Surveys Program. The reports are published in English and commonly a second language. The extracted tables were then translated from English by professional translators to all eight languages. The final dataset comprises 8,479 tables. For more information on the dataset creation, refer to the paper.

Visit

github.com

Connected records

paper

Tasks

data to textnatural language generation

Languages

HausaIgboSwahiliYoruba

Tags

Table-to-TextTaTa

Similaires

TaTa: A Multilingual Table-to-Text Dataset for African Languages

TaTa: A Multilingual Table-to-Text Dataset for African Languages

Existing data-to-text generation datasets are mostly limited to English. To address this lack of data, we create Table-to-Text in African languages (TaTa), the first large multilingual table-to-text dataset with a focus on African languages. We created TaTa by tran