Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

IDK-MRC: Unanswerable Questions for Indonesian Machine Reading Comprehension

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
PutOh,
Hôte:avatar
Machine Reading Comprehension (MRC) has become one of the essential tasks in Natural Language Understanding (NLU) as it is often included in several NLU benchmarks (Liang et al., 2020; Wilie et al., 2020). However, most MRC datasets only have answerable question type, overlooking the importance of unanswerable questions. MRC models trained only on answerable questions will select the span that is most likely to be the answer, even when the answer does not actually exist in the given passage (Rajpurkar et al., 2018). This problem especially remains in medium- to low-resource languages like Indonesian. Existing Indonesian MRC datasets (Purwarianti et al., 2007; Clark et al., 2020) are still inadequate because of the small size and limited question types, i.e., they only cover answerable questions. To fill this gap, we build a new Indonesian MRC dataset called I(n)don'tKnow- MRC (IDK-MRC) by combining the automatic and manual unanswerable question generation to minimize the cost of manual dataset construction while maintaining the dataset quality. Combined with the existing answerable questions, IDK-MRC consists of more than 10K questions in total. Our analysis shows that our dataset significantly improves the performance of Indonesian MRC models, showing a large improvement for unanswerable questions. EMNLP 2022

Visit

arxiv.org

Tasks

question answering

Tags

Computation and Language

Similaires

Question-Aware Deep Learning Model for Arabic Machine Reading ComprehensionY-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension with Open-Ended QuestionsA Swahili Question-Answering Dataset for Machine Reading Comprehension in HorticultureCan LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel DatasetEvaluating the Robustness of Machine Reading Comprehension Models to Low Resource Entity RenamingEDRAC: Benchmarking Arabic Dialect Reading Comprehension

Question-Aware Deep Learning Model for Arabic Machine Reading Comprehension

Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension with Open-Ended Questions

The purpose of this work is to share an English-Yorùbá evaluation dataset for openbook reading compr

A Swahili Question-Answering Dataset for Machine Reading Comprehension in Horticulture

Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset

Open-ended questions, which require students to produce multi-word, nontrivial responses, are a popu

Evaluating the Robustness of Machine Reading Comprehension Models to Low Resource Entity Renaming

Question answering (QA) models have shown compelling results in the task of Machine Reading Comprehe

EDRAC: Benchmarking Arabic Dialect Reading Comprehension

Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly