Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension with Open-Ended Questions

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AssAdeCheChu
Éditeur:
Und
Hôte:avatar
The purpose of this work is to share an English-Yorùbá evaluation dataset for openbook reading comprehension with open-ended questions to assess the performance of models both in a high- and a low-resource language. The dataset contains 358 questions and answers on 338 English documents and 208 Yorùbá documents. Experiments show a consistent disparity in performance between the two languages, with Yorùbá falling behind English for automatic metrics even if documents are much shorter for this language. For a small set of documents with comparable length, performance of Yorùbá drops by 2.5 times and this comparison is validated with human evaluation. When analyzing performance by length, we observe that Yorùbá decreases performance dramatically for documents that reach 1500 words while English performance is barely affected at that length. Our dataset opens the door to showcasing if English LLM reading comprehension capabilities extend to Yorùbá, which for the evaluated LLMs is not the case.

Visit

doi.orgunderline.io

Tasks

question answering

Languages

Yoruba

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similaires

A Realistic Rwandan Communty Health Worker Generated Vingette-based (I.e., Open-ended Questions) Benchmarking Dataset (with Associated Clinician and LLM Responses)Towards Open-Ended Discovery for Low-Resource NLPCan LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel DatasetIDK-MRC: Unanswerable Questions for Indonesian Machine Reading Comprehension2025/2026 Students’ Responses to Open-Ended Questions on COS101: Introduction to Computing at FUHSO, NigeriaStudents’ Responses to Open-Ended Questions on Computing Pioneers in COS101: Introduction to Computing at FUHSO, Nigeria

A Realistic Rwandan Communty Health Worker Generated Vingette-based (I.e., Open-ended Questions) Benchmarking Dataset (with Associated Clinician and LLM Responses)

The full dataset of 5,609 clinical vignettes contributed by 101 community health workers (CHWs) acro

Towards Open-Ended Discovery for Low-Resource NLP

Natural Language Processing (NLP) for low-resource languages remains fundamentally constrained by th

Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset

Open-ended questions, which require students to produce multi-word, nontrivial responses, are a popu

IDK-MRC: Unanswerable Questions for Indonesian Machine Reading Comprehension

Machine Reading Comprehension (MRC) has become one of the essential tasks in Natural Language Unders

2025/2026 Students’ Responses to Open-Ended Questions on COS101: Introduction to Computing at FUHSO, Nigeria

This study develops an AI-driven short-answer grading system that combines semantic sim

Students’ Responses to Open-Ended Questions on Computing Pioneers in COS101: Introduction to Computing at FUHSO, Nigeria

This dataset contains responses from 225 students enrolled in COS101: Introdu