Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Question-Answering in a Low-resourced Language: Benchmark Dataset and Models for Tigrinya

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
AssGaiParPar
Éditeur:
Und
Hôte:avatar
Question-Answering (QA) has seen significant advances recently, achieving near human-level performance over some benchmarks. However, these advances focus on high-resourced languages such as English, while the task remains unexplored for most other languages, mainly due to the lack of annotated datasets. This work presents a native QA dataset for an East African language, Tigrinya. The dataset contains 10.6K question-answer pairs spanning 572 paragraphs extracted from 290 news articles on various topics. The dataset construction method is discussed, which is applicable to constructing similar resources for related languages. We present comprehensive experiments and analyses of several resource-efficient approaches to QA, including monolingual, cross-lingual, and multilingual setups, along with comparisons against machine-translated silver data. Our strong baseline models reach 76% in the F1 score, while the estimated human performance is 92%, indicating that the benchmark presents a good challenge for future work. We make the dataset, models, and leaderboard publicly available.

Visit

doi.orgunderline.io

Tasks

question answering

Languages

Tigrigna

Tags

Natural Language Processing

Similaires

Tigrinya Question-Answering Dataset (TiQuAD)Context-Based Question Answering Using Large Language BERT Variant Models for Low Resourced Sesotho sa Leboa LanguageKenSwQuAD – A Question Answering Dataset for Swahili Low Resource LanguageTIGQA:An Expert Annotated Question Answering Dataset in TigrinyaAmaSQuAD: A Benchmark for Amharic Extractive Question AnsweringSwahiliVQA: A Dataset for Visual Question Answering in Swahili Language

Tigrinya Question-Answering Dataset (TiQuAD)

TiQuAD is a human-annotated question-answering dataset for the Tigrinya language. The dataset contai

Context-Based Question Answering Using Large Language BERT Variant Models for Low Resourced Sesotho sa Leboa Language

KenSwQuAD – A Question Answering Dataset for Swahili Low Resource Language

This research developed a Kencorpus Swahili Question Answering Dataset KenSwQuAD from raw data of Swahili language, which is a low resource language predominantly spoken in Eastern African and also has speakers in other parts of the world. Question Answering datase

TIGQA:An Expert Annotated Question Answering Dataset in Tigrinya

The absence of explicitly tailored, accessible annotated datasets for educational purposes presents

AmaSQuAD: A Benchmark for Amharic Extractive Question Answering

This research presents a novel framework for translating extractive question-answering datasets into

SwahiliVQA: A Dataset for Visual Question Answering in Swahili Language