Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BanHealthAIRespRelv3: A Benchmark Dataset for Evaluating Bengali AI-Generated Health Responses Based on User-Specific Prompts

Domaine:

natural language processinghealthcare

Type de record:

dataset
Créateur:
ZahTanBisBij
Éditeur:
Daffodil International University
Éditeur:
Men
Hôte:avatar
BanHealthAIRespRelv3 serves as a benchmark dataset for evaluating the relevance and safety of AI-generated health responses in Bengali, a low-resource language. Its primary goal is to address risks of medical hallucinations, irrelevant advice, and misinformation in online health information seeking (OHIS), particularly in resource-constrained settings like Bangladesh. To ensure high quality, the dataset underwent rigorous preprocessing (text cleaning, duplicate removal, filtering incomplete responses) and annotation by three native Bengali speakers, with reliability confirmed via Fleiss’ Kappa (substantial to almost perfect agreement) and retention of only high-confidence instances (>0.8). BanHealthAIRespRelv3 is a ternary-class dataset comprising 15,384 instances of user-simulated health prompts and corresponding AI responses (from ChatGPT and Google Gemini), categorized as: Highly Relevant: 5,004 instances Partially Relevant: 5,300 instances Not Relevant: 5,084 instances The dataset has broad applications in multiple NLP and AI areas, including: Relevance and hallucination detection in health AI Development of safer, culturally sensitive health chatbots Trust calibration and ethical AI in digital health Misinformation mitigation in low-resource languages Educational and research advancements in Bengali medical NLP BanHealthAIRespRelv3 is openly available for academic and research purposes, promoting collaboration and innovation in Bengali NLP. By establishing a robust benchmark, it aims to foster trustworthy, context-aware AI systems for health-related applications in under-resourced languages.

Visit

doi.orgdata.mendeley.com

Tasks

text classification

Tags

Natural Language ProcessingHealthEthical LLM

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

KrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural AdvisoryBenchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word EmbeddingsNigerian Academic Writing Corpus: Pre-AI Benchmark for AI-Generated Text DetectionALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text DetectionA Realistic Rwandan Communty Health Worker Generated Vingette-based (I.e., Open-ended Questions) Benchmarking Dataset (with Associated Clinician and LLM Responses)AfriEconQA: A Benchmark Dataset for African Economic Analysis based on World Bank Reports

KrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural Advisory

We present KrishokChat, the first citation-grounded Bengali agricultural instruction-tuning dataset

Benchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embeddings. DiaLex covers five importa

Nigerian Academic Writing Corpus: Pre-AI Benchmark for AI-Generated Text Detection

A curated corpus of pre-AI era (2005–2022) Nigerian academic writing paired with AI-generated equiva

ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection

We introduce ALHD, the first large-scale comprehensive Arabic dataset explicitly designed to disting

A Realistic Rwandan Communty Health Worker Generated Vingette-based (I.e., Open-ended Questions) Benchmarking Dataset (with Associated Clinician and LLM Responses)

The full dataset of 5,609 clinical vignettes contributed by 101 community health workers (CHWs) acro

AfriEconQA: A Benchmark Dataset for African Economic Analysis based on World Bank Reports

We introduce AfriEconQA, a specialized benchmark dataset for African economic analysis grounded in a