Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TaTA: A Multilingual Table-to-Text Dataset for African Languages

Domain:

natural language processinghealthcare

Record type:

dataset

TaTA (Table-to-Text in African languages) is the first large multilingual table-to-text datasets with a focus on African languages. The dataset is parallel and covers nine languages, eight of which are spoken in Africa: Arabic, English, French, Hausa, Igbo, Portuguese, Swahili, Yorùbá, and Russian.

TaTa was created by extracting tables from charts in 71 PDF reports published between 1990 and 2021 by the Demographic and Health Surveys Program. The reports are published in English and commonly a second language. The extracted tables were then translated from English by professional translators to all eight languages. The final dataset comprises 8,479 tables. For more information on the dataset creation, refer to the paper.

Visit

github.com

Connected records

paper

Tasks

data to textnatural language generation

Languages

HausaIgboSwahiliYoruba

Tags

Table-to-TextTaTa

Similar

TaTa: A Multilingual Table-to-Text Dataset for African Languages

TaTa: A Multilingual Table-to-Text Dataset for African Languages

Existing data-to-text generation datasets are mostly limited to English. To address this lack of data, we create Table-to-Text in African languages (TaTa), the first large multilingual table-to-text dataset with a focus on African languages. We created TaTa by tran