Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CPIQA: Climate Paper Image Question Answering Dataset for Retrieval-Augmented Generation with Context-based Query Expansion

Domain:

natural language processingclimate

Record type:

dataset
Creator:
MutMid
Publisher:
Zenodo
Host:avatar
CPIQA is a large scale QA dataset focused on figured extracted from scientific research papers from various peer reviewed venues in the climate science domain. The figures extracted include tables, graphs and diagrams, which inform the generation of questions using large language models (LLMs). Notably this dataset includes questions for 3 audiences: general public, climate skeptic and climate expert. 4 types of questions are generated with various focusses including figures, numerical, text-only and general. This results in 12 questions generated per scientific paper. Alongside figures, descriptions of the figures generated using multimodal LLMs are included and used. This work was funded through the WCSSP South Africa project, a collaborative initiative between the Met Office, South African and UK partners, supported by the International Science Partnership Fund (ISPF) from the UK's Department for Science, Innovation and Technology (DSIT). It is also supported by the Natural Environment Research Council (grant NE/S015604/1) project GloSAT.Mutalik, R. Panchalingam, A. Loitongbam, G. Osborn, T. J. Hawkins, E. Middleton, S. E. CPIQA: Climate Paper Image Question Answering Dataset for Retrieval-Augmented Generation with Context-based Query Expansion, ClimateNLP-2025, ACL, 31st July 2025, nlp4climate.github.io

Visit

doi.orgzenodo.org

Tasks

question answering

Tags

Machine LearningNatural Language ProcessingArtificial IntelligenceEnvironmental Science

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Improving Amharic Legal Question Answering with Retrieval-Augmented Generation and Locally-Sourced DataBi-gram based Query Expansion Technique for Amharic Information Retrieval SystemQuery Expansion Based-on Similarity of Terms for Improving Arabic Information RetrievalSwahili Question-Answering Dataset for HorticultureLuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question AnsweringAn ontology driven question answering system for fatawa retrieval

Improving Amharic Legal Question Answering with Retrieval-Augmented Generation and Locally-Sourced Data

Bi-gram based Query Expansion Technique for Amharic Information Retrieval System

Query Expansion Based-on Similarity of Terms for Improving Arabic Information Retrieval

Part 6: Information Retrieval International audience This research suggests a method

Swahili Question-Answering Dataset for Horticulture

The dataset was created to contribute to Swahili language resources for natural language processing

LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully rec

An ontology driven question answering system for fatawa retrieval

This work aims to propose a system for the Algerian Fatawa House in orderto facilitate the task of t