Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

Domain:

healthcarenatural language processing

Record type:

paper
Creator:
KimWu,KimVas
Publisher:
arXiv
Host:avatar
Recent advancements in multimodal Large Language Models (LLMs) have significantly enhanced the automation of medical image analysis, particularly in generating radiology reports from chest X-rays (CXR). However, these models still suffer from hallucinations and clinically significant errors, limiting their reliability in real-world applications. In this study, we propose Look & Mark (L&M), a novel grounding fixation strategy that integrates radiologist eye fixations (Look) and bounding box annotations (Mark) into the LLM prompting framework. Unlike conventional fine-tuning, L&M leverages in-context learning to achieve substantial performance gains without retraining. When evaluated across multiple domain-specific and general-purpose models, L&M demonstrates significant gains, including a 1.2% improvement in overall metrics (A.AVG) for CXR-LLaVA compared to baseline prompting and a remarkable 9.2% boost for LLaVA-Med. General-purpose models also benefit from L&M combined with in-context learning, with LLaVA-OV achieving an 87.3% clinical average performance (C.AVG)-the highest among all models, even surpassing those explicitly trained for CXR report generation. Expert evaluations further confirm that L&M reduces clinically significant errors (by 0.43 average errors per report), such as false predictions and omissions, enhancing both accuracy and reliability. These findings highlight L&M's potential as a scalable and efficient solution for AI-assisted radiology, paving the way for improved diagnostic workflows in low-resource clinical settings.

Visit

doi.orgarxiv.org

Tasks

computer visionimage-text retrieval

Tags

Computer Vision and Pattern Recognition (cs.CV)Computation and Language (cs.CL)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

PaliGemma-CXR: a multi-task multimodal model for tuberculosis chest X-ray interpretationNigeria Chest X-ray Datasetasmelashteka/Afro-Chest-X-rayConventional and Lodox chest x-ray imageshaftomdesbele/TRANSFER-LEARNING-AND-KNOWLEDGE-DISTILLATION-for-CHEST-X-RAY-CLASSIFICATIONSchema Generation for Large Knowledge Graphs Using Large Language Models

PaliGemma-CXR: a multi-task multimodal model for tuberculosis chest X-ray interpretation

Abstract Background Uganda has a hig

Nigeria Chest X-ray Dataset

asmelashteka/Afro-Chest-X-ray

Chest X-ray Imaging Dataset for Multiple Cardio-respiratory Diseases in Ethiopia # Afro Chest X-ray

Conventional and Lodox chest x-ray images

This is a dataset used in prospective and retrospective designs to compare conventional x-ray system

haftomdesbele/TRANSFER-LEARNING-AND-KNOWLEDGE-DISTILLATION-for-CHEST-X-RAY-CLASSIFICATION

TRANSFER LEARNING AND KNOWLEDGE DISTILLATION for CHEST X-RAY CLASSIFICATION, This makes the approach

Schema Generation for Large Knowledge Graphs Using Large Language Models

Schemas play a vital role in ensuring data quality and supporting usability in the Semantic Web and