paper-with-me

Papers

ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation

2026-02-04 · Jiarui Jin, Haoyu Wang, Xingliang Wu, Xiaocheng Fang, Xiang Lan, Zihan Wang, Deyun Zhang, Bo Liu, Yingying Zhang, Xian Wu, Hongyan Li, Shenda Hong arxiv

Electrocardiography (ECG) serves as an indispensable diagnostic tool in clinical practice, yet existing multimodal large language models (MLLMs) remain unreliable for ECG interpretation, often producing plausible but clinically incorrect analyses. To address this, we propose ECG-R1, the first reasoning ECG MLLM designed for reliable ECG interpretation via three innovations. First, we construct the interpretation corpus using \textit{Protocol-Guided Instruction Data Generation}, grounding interpretation in measurable ECG features and monograph-defined quantitative thresholds and diagnostic logic. Second, we present a modality-decoupled architecture with \textit{Interleaved Modality Dropout} to improve robustness and cross-modal consistency when either the ECG signal or ECG image is missing. Third, we present \textit{Reinforcement Learning with ECG Diagnostic Evidence Rewards} to strengthen evidence-grounded ECG interpretation. Additionally, we systematically evaluate the ECG interpretation capabilities of proprietary, open-source, and medical MLLMs, and provide the first quantitative evidence that severe hallucinations are widespread, suggesting that the public should not directly trust these outputs without independent verification. Code is available at \href{https://github.com/PKUDigitalHealth/ECG-R1}{here}.

📄 PDF Abstract BibTeX arXiv:2602.04279

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing

2026-07-09 · Dominick Reilly, Qiyu Wu, Hiromi Wakaki, Srijan Das 외 arxiv

Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, many real-world settings violate this assu…

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

2026-08-05 · Yue Zhang, Yingzhao Jian, Yunqiu Xu, Xiaoxiao Sun 외 hf

Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these mod…

Multimodal ReasoningScene Understanding

Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent

2026-05-10 · Yihong Tang, Kehai Chen, Xuefeng Bai, Min Zhang arxiv

The advancement of Multimodal Large Language Models (MLLMs) has expanded Role-Playing Agents (RPAs) into visually grounded environments. However, human vision is inherently subjective and identity-driven, whereas existin…

Visual Grounding

PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

2026-09-04 · Yuchen Sun, Qian Yang, Jun Wang, Detai Xin 외 arxiv

Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio modality. Existing benchmarks either treat audio as an auxiliary component of video quality or assess it in is…

Audio GenerationVideo Generation

PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection

2025-10-27 · Yusu Qian, Cheng Wan, Chao Jia, Yinfei Yang 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable progress on vision-language tasks, yet their reasoning processes remain sometimes unreliable. We introduce PRISM-Bench, a benchmark of puzzle-based visual…

Multimodal ReasoningAnswer GenerationVisual Reasoning