Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases

Domaine:

healthcare

Type de record:

paperdataset
Créateur:
Chen, QiLi,BasZho
Hôte:avatar
Artificial intelligence (AI) has achieved remarkable success in medical imaging, but it is widely recognized that these models often perform inconsistently across real-world clinical settings. Such inconsistencies occur when patient demographics and imaging protocols vary, for example, in detecting small tumors, analyzing scans from different contrast phases, or evaluating patients of different ages or sexes. To quantify these inconsistencies, we develop a large-scale, open benchmark of 85,355 CT scans that systematically evaluates 12 tumor-detection AI models across tumor size, location, patient subgroup, and imaging protocol. We leverage large language models (LLMs) to extract and organize subgroup information from clinical data, which makes the analysis both scalable and reproducible. Our benchmark reveals that current state-of-the-art AI models, optimized for average accuracy, perform poorly in rare or underrepresented subgroups, such as young, female African Americans. However, collecting sufficient annotated data for these rare cases is often impractical. The benchmark provides a foundation for building more reliable and robust AI models for tumor detection and highlighting the need for rigorous, subgroup-level evaluation in medical imaging and computer vision. Datasets, code

Visit

arxiv.org

Tasks

computer visionimage classification

Tags

Computer Vision and Pattern Recognition

Similaires

Participatory AI for Agriculture: Co-Designing Hydroponic Pest Detection Models with African FarmersThe African Breast Imaging Dataset for Equitable Cancer Care: Protocol for an Open Mammogram and Ultrasound Breast Cancer Detection DatasetNGLUEni: Benchmarking and Adapting Pretrained Language Models for Nguni Languagesshiramichel/ai-speech-biasesAfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African LanguagesBenchmarking Vision Language Models for Cultural Understanding

Participatory AI for Agriculture: Co-Designing Hydroponic Pest Detection Models with African Farmers

Africa faces pressing agricultural challenges exacerbated by climate change, declining productivity,

The African Breast Imaging Dataset for Equitable Cancer Care: Protocol for an Open Mammogram and Ultrasound Breast Cancer Detection Dataset

Abstract Introduction Breast cancer is one

NGLUEni: Benchmarking and Adapting Pretrained Language Models for Nguni Languages

shiramichel/ai-speech-biases

English-language stock AI synthetic voices and voice clones generated using Speechify and ElevenLabs

AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages

Text embeddings are an essential building component of several NLP tasks such as retrieval-augmented

Benchmarking Vision Language Models for Cultural Understanding

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLM