paper-with-me

홈 › Papers

RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze

2025-07-12 · Yunsoo Kim, Jinge Wu, Honghan Wu arxiv

Large Vision-Language Models (LVLMs) have demonstrated promising performance in chest X-ray (CXR) analysis. To enhance human-computer interaction, several studies have incorporated radiologists' eye gaze, typically through heatmaps or textual prompts. However, these methods often overlook the sequential order of eye movements, which could provide valuable insights by highlighting both the areas of interest and the order in which they are examined. In this work, we propose a novel approach called RadEyeVideo that integrates radiologists' eye-fixation data as a video sequence, capturing both the temporal and spatial dynamics of their gaze. We evaluate this method in CXR report generation and disease diagnosis using three general-domain, open-source LVLMs with video input capabilities. When prompted with eye-gaze videos, model performance improves by up to 24.6% in the report generation task and on average 15.2% for both tasks using scaled evaluation metrics. Notably, RadEyeVideo enhanced an open-domain LVLM model, LLaVA-OneVision, to surpass task-specific medical LVLMs such as MAIRA-2 and CheXagent, trained on large Chest X-ray data. This work highlights that domain expert's knowledge (eye-gaze information in this case), when effectively integrated with LVLMs, can significantly enhance general-domain models' capabilities in clinical tasks. RadEyeVideo is a step toward a scalable human-centered approach of utilizing LVLMs in medical image analytics.

📄 PDF Abstract BibTeX arXiv:2507.09097

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Robust Domain Generalization for Multi-modal Object Recognition

2024-08-11 · Yuxin Qiao, Keqin Li, Junhong Lin, Rong Wei 외

In multi-label classification, machine learning encounters the challenge of domain generalization when handling tasks with distributions differing from the training data. Existing approaches primarily focus on vision obj…

Domain GeneralizationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONObject+1

Cross-Domain Foundation Model Adaptation: Pioneering Computer Vision Models for Geophysical Data Analysis

2024-08-22 · Zhixiang Guo, Xinming Wu, Luming Liang, Hanlin Sheng 외

We explore adapting foundation models (FMs) from the computer vision domain to geoscience. FMs, large neural networks trained on massive datasets, excel in diverse tasks with remarkable adaptability and generality. Howev…

Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection

2025-02-18 · Jingbiao Mei, Jinghong Chen, Guangyu Yang, Weizhe Lin 외

Hateful memes have become a significant concern on the Internet, necessitating robust automated detection systems. While LMMs have shown promise in hateful meme detection, they face notable challenges like sub-optimal pe…

Contrastive LearningDomain GeneralizationHateful Meme ClassificationIn-Context Learning+2

A Foundation Language-Image Model of the Retina (FLAIR): Encoding Expert Knowledge in Text Supervision

2023-08-15 · Julio Silva-Rodríguez, Hadi Chakor, Riadh Kobbi, Jose Dolz 외

Foundation vision-language models are currently transforming computer vision, and are on the rise in medical imaging fueled by their very promising generalization capabilities. However, the initial attempts to transfer t…

DescriptiveLanguage Modelling

Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction

2025-04-27 · Xiaoran Xu, Jiangang Yang, Wenyue Chong, Wenhui Shi 외

Single-Domain Generalized Object Detection~(S-DGOD) aims to train an object detector on a single source domain while generalizing well to diverse unseen target domains, making it suitable for multimedia applications that…

object-detectionObject Detection