Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Multimodal Attention-Based Multi-Instance Learning Framework for Fair and Interpretable Pediatric Teledermatology

Domain:

healthcare

Record type:

papermodel
Creator:
Kho
Publisher:
Spr
Host:
Abstract Purpose : Pediatric skin diseases are prevalent yet frequently underdiagnosed in low-resource settings across Sub-Saharan Africa due to limited access to specialized dermatological care. This study examines whether a subject-level multimodal learning framework can improve diagnostic accuracy, interpretability, and fairness in pediatric teledermatology across diverse skin types. Methods : A subject-level multimodal multi-instance learning framework is developed in which each patient is represented as a bag of clinical images, with visual features integrated alongside demographic and clinical metadata. A gated attention mechanism is employed to aggregate heterogeneous image instances into interpretable subject-level representations, while multimodal fusion provides contextual information for diagnosis. The framework is evaluated using the PASSION pediatric dermatology dataset across four common skin conditions. Ablation studies and statistical analyses are conducted to assess the contributions of attention-based aggregation and multimodal fusion. Fairness is evaluated across Fitzpatrick skin types. Results : The proposed framework achieves an overall classification accuracy of 82.8\% and a macro F1-score of 0.81. Ablation results demonstrate that gated attention-based aggregation significantly outperforms naive pooling strategies, while multimodal fusion further enhances diagnostic robustness. Fairness analysis indicates stable performance across Fitzpatrick skin types. Conclusion : Subject-level multimodal learning provides a robust, interpretable, and equitable approach for AI-assisted pediatric teledermatology, demonstrating strong potential for improving diagnostic access and quality of care in low-resource clinical environments.

Visit

doi.org

Tasks

computer visionimage classification

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

MAIN: Multi-Attention Instance Network for Video SegmentationMulti-Head Attention with Diversity for Learning Grounded Multilingual Multimodal RepresentationsMulti-stream Attention-Enhanced Deep Learning Framework for Cocoa Leaf Disease Detection and Classification in GhanaMultimodal based Amharic fake news detection using CNN and attention-based BiLSTMAttention Based Hybrid Deep Learning models for Multi class Amharic News Categorization with Explainable AIRetrieval Augmented Enhanced Dual Co-Attention Framework for Target Aware Multimodal Bengali Hateful Meme Detection

MAIN: Multi-Attention Instance Network for Video Segmentation

Instance-level video segmentation requires a solid integration of spatial and temporal information.

Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations

With the aim of promoting and understanding the multilingual version of image search, we leverage vi

Multi-stream Attention-Enhanced Deep Learning Framework for Cocoa Leaf Disease Detection and Classification in Ghana

Multimodal based Amharic fake news detection using CNN and attention-based BiLSTM

Attention Based Hybrid Deep Learning models for Multi class Amharic News Categorization with Explainable AI

Abstract Efficient and adaptable text classification systems are required due to the incre

Retrieval Augmented Enhanced Dual Co-Attention Framework for Target Aware Multimodal Bengali Hateful Meme Detection

Hateful content on social media increasingly appears as multimodal memes that combine images and tex