paper-with-me

홈 › Papers

CheXPO-v2: Preference Optimization for Chest X-ray VLMs with Knowledge Graph Consistency

2025-12-19 · Xiao Liang, Yuxuan An, Di Wang, Jiawei Hu, Zhicheng Jiao, Bin Jing, Quan Wang arxiv

Medical Vision-Language Models (VLMs) are prone to hallucinations, compromising clinical reliability. While reinforcement learning methods like Group Relative Policy Optimization (GRPO) offer a low-cost alignment solution, their reliance on sparse, outcome-based rewards inadvertently encourages models to "overthink" -- generating verbose, convoluted, and unverifiable Chain-of-Thought reasoning to justify answers. This focus on outcomes obscures factual errors and poses significant safety risks. To address this, we propose CheXPO-v2, a novel alignment framework that shifts from outcome to process supervision. Our core innovation is a Knowledge Graph Consistency Reward mechanism driven by Entity-Relation Matching. By explicitly parsing reasoning steps into structured "Disease, Relation, Anatomy" triplets, we provide fine-grained supervision that penalizes incoherent logic and hallucinations at the atomic level. Integrating this with a hard-example mining strategy, our approach significantly outperforms GRPO and state-of-the-art models on benchmarks like MIMIC-CXR-VQA. Crucially, CheXPO-v2 achieves new state-of-the-art accuracy using only 5k samples, demonstrating exceptional data efficiency while producing clinically sound and verifiable reasoning. The project source code is publicly available at: https://github.com/ecoxial2007/CheX-Phi4MM.

📄 PDF Abstract BibTeX arXiv:2512.17213

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale

2025-07-09 · Xiao Liang, Jiawei Hu, Di Wang, Zhi Ma 외 arxiv

Vision-language models (VLMs) are prone to hallucinations that critically compromise reliability in medical applications. While preference optimization can mitigate these hallucinations through clinical feedback, its imp…

Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation

2024-06-10 · Oishi Banerjee, Hong-Yu Zhou, Subathra Adithan, Stephen Kwak 외

Recent advances in generative vision-language models (VLMs) have exciting potential implications for AI in radiology, yet VLMs are also known to produce hallucinations, nonsensical text, and other unwanted behaviors that…

Preference Fine-Tuning for Factuality in Chest X-Ray Interpretation Models Without Human Feedback

2024-10-09 · Dennis Hein, Zhihong Chen, Sophie Ostmeier, Justin Xu 외

Radiologists play a crucial role by translating medical images into medical reports. However, the field faces staffing shortages and increasing workloads. While automated approaches using vision-language models (VLMs) sh…

MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

2024-12-09 · Kangyu Zhu, Peng Xia, Yun Li, Hongtu Zhu 외

The advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) encounter factuality challenges due to modality misalignment, where the mod…

Visual Question Answering (VQA)

Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards

2025-03-25 · Alexander Gambashidze, Konstantin Sobolev, Andrey Kuznetsov, Ivan Oseledets

Can Visual Language Models (VLMs) effectively capture human visual preferences? This work addresses this question by training VLMs to think about preferences at test time, employing reinforcement learning methods inspire…

World Knowledge