Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages

Domain:

natural language processing

Record type:

paper
Confidently making progress on multilingual modeling requires challenging, trustworthy evaluations. We present TyDi QA---a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages of TyDi QA are diverse with regard to their typology---the set of linguistic features each language expresses---such that we expect models performing well on this set to generalize across a large number of the world's languages. We present a quantitative analysis of the data quality and example-level qualitative linguistic analyses of observed language phenomena that would not be found in English-only corpora. To provide a realistic information-seeking task and avoid priming effects, questions are written by people who want to know the answer, but don't know the answer yet, and the data is collected directly in each language without the use of translation.

Visit

arxiv.orgaclanthology.org

Connected records

dataset

Tasks

question answering

Languages

Swahili

Licenses

Similar

TyDi QA-WANA: A Benchmark for Information-Seeking Question Answering in Languages of West Asia and North AfricaAfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark DatasetAmaSQuAD: A Benchmark for Amharic Extractive Question AnsweringUhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages

TyDi QA-WANA: A Benchmark for Information-Seeking Question Answering in Languages of West Asia and North Africa

We present TyDi QA-WANA, a question-answering dataset consisting of 28K examples divided among 10 la

AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset

Recent advancements in large language model(LLM) performance on medical multiple choice question (MC

AmaSQuAD: A Benchmark for Amharic Extractive Question Answering

This research presents a novel framework for translating extractive question-answering datasets into

Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages

Evaluations of Large Language Models (LLMs) on knowledge-intensive tasks and factual accuracy often