paper-with-me

Papers

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

2026-07-30 · Zhilin Wu, Zhangkai Ni, Chengmei Yang, Longzhen Yang, Yihang Liu, Ying Wen, Lianghua He arxiv

In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal information is crucial for reliable assessment of disease progression and treatment response. However, despite the rapid advancement of multimodal large language models (MLLMs), longitudinal medical visual reasoning remains largely underexplored. To fill this gap, we propose LoMeVQA, a comprehensive benchmark consisting of 206K longitudinal visual question answering (VQA) pairs for temporal medical image analysis. LoMeVQA covers five tasks: progress classification, progress description, progress report generation, differential region grounding, and differential region description. To construct the dataset, we develop an automated pipeline that (1) organizes patient records chronologically, (2) extracts clinically meaningful entities via a medical knowledge graph, and (3) models their temporal evolution to guide large language models in generating high-quality longitudinal VQA pairs. Extensive evaluations demonstrate that both general-purpose and medical-domain MLLMs perform poorly on LoMeVQA, revealing substantial limitations in temporal reasoning. To address these limitations, we introduce MedLong-8B, which achieves state-of-the-art performance across all tasks. Beyond benchmarking, we conduct detailed analyses that uncover key failure modes and shed light on how to improve longitudinal medical visual reasoning. Our data is available at: https://github.com/pepperbubble/LoMeVQA

📄 PDF Abstract BibTeX arXiv:2607.27806

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringVisual Reasoning

Similar Papers 제목 키워드 기반

TIMER: Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records

2025-03-06 · Hejie Cui, Alyssa Unell, Bowen Chen, Jason Alan Fries 외

Large language models (LLMs) have emerged as promising tools for assisting in medical tasks, yet processing Electronic Health Records (EHRs) presents unique challenges due to their longitudinal nature. While LLMs' capabi…

LesionLocator: Zero-Shot Universal Tumor Segmentation and Tracking in 3D Whole-Body Imaging

2025-01-01 · CVPR 2025 1 · Maximilian Rokuss, Yannick Kirchhoff, Seval Akbal, Balint Kovacs 외

In this work, we present LesionLocator, a framework for zero-shot longitudinal lesion tracking and segmentation in 3D medical imaging, establishing the first end-to-end model capable of 4D tracking with dense spatial…

Lesion SegmentationSegmentationTumor Segmentation

MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays

2026-05-15 · Sunghwan Steve Cho, Yunseok Han, Jaeyoung Do arxiv

Longitudinal chest X-ray (CXR) interpretation requires reasoning over disease evolution across multiple patient visits, yet most existing medical VQA benchmarks focus on single images or short-horizon image pairs. We int…

Visual Reasoning

HERGen: Elevating Radiology Report Generation with Longitudinal Data

2024-07-21 · Fuying Wang, Shenghui Du, Lequan Yu

Radiology reports provide detailed descriptions of medical imaging integrated with patients' medical histories, while report writing is traditionally labor-intensive, increasing radiologists' workload and the risk of dia…

Diagnostic

SHAPE: A Sample-adaptive Hierarchical Prediction Network for Medication Recommendation

2023-09-09 · Sicen Liu, Xiaolong Wang, Jingcheng Du, Yongshuai Hou 외

Effectively medication recommendation with complex multimorbidity conditions is a critical task in healthcare. Most existing works predicted medications based on longitudinal records, which assumed the information transm…