Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Question-Answering in a Low-resourced Language: Benchmark Dataset and Models for Tigrinya

Domain:

natural language processing

Record type:

datasetpaper
Creator:
AssGaiParPar
Publisher:
Und
Host:avatar
Question-Answering (QA) has seen significant advances recently, achieving near human-level performance over some benchmarks. However, these advances focus on high-resourced languages such as English, while the task remains unexplored for most other languages, mainly due to the lack of annotated datasets. This work presents a native QA dataset for an East African language, Tigrinya. The dataset contains 10.6K question-answer pairs spanning 572 paragraphs extracted from 290 news articles on various topics. The dataset construction method is discussed, which is applicable to constructing similar resources for related languages. We present comprehensive experiments and analyses of several resource-efficient approaches to QA, including monolingual, cross-lingual, and multilingual setups, along with comparisons against machine-translated silver data. Our strong baseline models reach 76% in the F1 score, while the estimated human performance is 92%, indicating that the benchmark presents a good challenge for future work. We make the dataset, models, and leaderboard publicly available.

Visit

doi.orgunderline.io

Tasks

question answering

Languages

Tigrigna

Tags

Natural Language Processing

Similar

Tigrinya Question-Answering Dataset (TiQuAD)Context-Based Question Answering Using Large Language BERT Variant Models for Low Resourced Sesotho sa Leboa LanguageKenSwQuAD – A Question Answering Dataset for Swahili Low Resource LanguageTIGQA:An Expert Annotated Question Answering Dataset in TigrinyaAmaSQuAD: A Benchmark for Amharic Extractive Question AnsweringSwahiliVQA: A Dataset for Visual Question Answering in Swahili Language

Tigrinya Question-Answering Dataset (TiQuAD)

TiQuAD is a human-annotated question-answering dataset for the Tigrinya language. The dataset contai

Context-Based Question Answering Using Large Language BERT Variant Models for Low Resourced Sesotho sa Leboa Language

KenSwQuAD – A Question Answering Dataset for Swahili Low Resource Language

This research developed a Kencorpus Swahili Question Answering Dataset KenSwQuAD from raw data of Swahili language, which is a low resource language predominantly spoken in Eastern African and also has speakers in other parts of the world. Question Answering datase

TIGQA:An Expert Annotated Question Answering Dataset in Tigrinya

The absence of explicitly tailored, accessible annotated datasets for educational purposes presents

AmaSQuAD: A Benchmark for Amharic Extractive Question Answering

This research presents a novel framework for translating extractive question-answering datasets into

SwahiliVQA: A Dataset for Visual Question Answering in Swahili Language