Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

DREAM: A Challenge Dataset and Models for Dialogue-Based Reading Comprehension

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
SunYu,CheYu,
Hôte:avatar
We present DREAM, the first dialogue-based multiple-choice reading comprehension dataset. Collected from English-as-a-foreign-language examinations designed by human experts to evaluate the comprehension level of Chinese learners of English, our dataset contains 10,197 multiple-choice questions for 6,444 dialogues. In contrast to existing reading comprehension datasets, DREAM is the first to focus on in-depth multi-turn multi-party dialogue understanding. DREAM is likely to present significant challenges for existing reading comprehension systems: 84% of answers are non-extractive, 85% of questions require reasoning beyond a single sentence, and 34% of questions also involve commonsense knowledge. We apply several popular neural reading comprehension models that primarily exploit surface information within the text and find them to, at best, just barely outperform a rule-based approach. We next investigate the effects of incorporating dialogue structure and different kinds of general world knowledge into both rule-based and (neural and non-neural) machine learning-based reading comprehension models. Experimental results on the DREAM dataset show the effectiveness of dialogue structure and general world knowledge. DREAM will be available at dataset.org. To appear in TACL

Visit

arxiv.org

Tasks

commonsense reasoningquestion answering

Tags

Computation and Language

Similaires

NaijaRC: A Multi-choice Reading Comprehension Dataset for Nigerian LanguagesA Swahili Question-Answering Dataset for Machine Reading Comprehension in HorticultureThe Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsBreaking the cycle of poor critical reading comprehension: A strategy-based interventionAccelerating cough-based algorithms for pulmonary tuberculosis screening: Results from the CODA TB DREAM ChallengeY-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension with Open-Ended Questions

NaijaRC: A Multi-choice Reading Comprehension Dataset for Nigerian Languages

In this paper, we create NaijaRC: a new multi-choice Reading Comprehension dataset for three native

A Swahili Question-Answering Dataset for Machine Reading Comprehension in Horticulture

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

We present Belebele, a multiple-choice machine reading comprehension (MRC) dataset spanning 122 lang

Breaking the cycle of poor critical reading comprehension: A strategy-based intervention

Critical reading comprehension is a crucial skill for lifelong learning, but many learners struggle

Accelerating cough-based algorithms for pulmonary tuberculosis screening: Results from the CODA TB DREAM Challenge

Abstract Importance Open-access data chall

Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension with Open-Ended Questions

The purpose of this work is to share an English-Yorùbá evaluation dataset for openbook reading compr