Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriEconQA: A Benchmark Dataset for African Economic Analysis based on World Bank Reports

Domain:

natural language processingsocioeconomic

Record type:

paperdataset
Creator:
Aja
Host:avatar
We introduce AfriEconQA, a specialized benchmark dataset for African economic analysis grounded in a comprehensive corpus of 236 World Bank reports. The task of AfriEconQA is to answer complex economic queries that require high-precision numerical reasoning and temporal disambiguation from specialized institutional documents. The dataset consists of 8,937 curated QA instances, rigorously filtered from a pool of 10018 synthetic questions to ensure high-quality evidence-answer alignment. Each instance is composed of: (1) a question requiring reasoning over economic indicators, (2) the corresponding evidence retrieved from the corpus, (3) a verified ground-truth answer, and (4) source metadata (e.g., URL and publication date) to ensure temporal provenance. AfriEconQA is the first benchmark focused specifically on African economic analysis, providing a unique challenge for Information Retrieval (IR) systems, as the data is largely absent from the pretraining corpora of current Large Language Models (LLMs). We operationalize this dataset through an 11-experiment matrix, benchmarking a zero-shot baseline (GPT-5 Mini) against RAG configurations using GPT-4o and Qwen 32B across five distinct embedding and ranking strategies. Our results demonstrate a severe parametric knowledge gap, where zero-shot models fail to answer over 90 percent of queries, and even state-of-the-art RAG pipelines struggle to achieve high precision. This confirms AfriEconQA as a robust and challenging benchmark for the next generation of domain-specific IR and RAG systems. The AfriEconQA dataset and code will be made publicly available upon publication.

Visit

arxiv.org

Tasks

question answering

Tags

Computation and Language

Similar

HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African LanguageTransformer-Based Clinical Annotation of Lung Cancer Reports: A Benchmark and Fine-Tuning Study on a Novel Tunisian CorpusBanHealthAIRespRelv3: A Benchmark Dataset for Evaluating Bengali AI-Generated Health Responses Based on User-Specific PromptsFields of The World: A Machine Learning Benchmark Dataset For Global Agricultural Field Boundary SegmentationAfriSenti: A Twitter Sentiment Analysis Benchmark for African LanguagesBenchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African Language

The development of Natural Language Processing (NLP) tools for low-resource languages is critically

Transformer-Based Clinical Annotation of Lung Cancer Reports: A Benchmark and Fine-Tuning Study on a Novel Tunisian Corpus

Background: Lung cancer causes more deaths than any other malignancy worldwide, accounting for 2.2 m

BanHealthAIRespRelv3: A Benchmark Dataset for Evaluating Bengali AI-Generated Health Responses Based on User-Specific Prompts

BanHealthAIRespRelv3 serves as a benchmark dataset for evaluating the relevance and safety of AI-gen

Fields of The World: A Machine Learning Benchmark Dataset For Global Agricultural Field Boundary Segmentation

Crop field boundaries are foundational datasets for agricultural monitoring and assessments but are

AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages

Africa is home to over 2000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages

Benchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embeddings. DiaLex covers five importa