paper-with-me

홈 › Papers

Model Merging for Medical LVLMs: A Benchmark and a Winner-Take-All Approach

2026-07-17 · Lichao Mou, Shilan Zhang, Chunlei Li, Bingcong Yan, Jingliang Hu, Yilei Shi, Shengwu Xiong, Xiao Xiang Zhu, Lei Li, Yaxiong Chen arxiv

Large vision-language models (LVLMs) can be adapted to specialized medical imaging tasks via parameter-efficient fine-tuning approaches such as low-rank adaptation (LoRA), leading to a growing ecosystem of expert models tailored to specific imaging modalities and clinical scenarios. However, deploying multiple expert LVLMs in practice incurs substantial computational and operational overhead. Model merging provides a promising solution by consolidating multiple experts into a single model without retraining, yet it remains largely unexplored in the medical domain. In this work, we present the first systematic study of model merging for medical LVLMs. We introduce MergeMedBench, a comprehensive benchmark spanning eight imaging modalities and diverse clinical task types, comprising 16 LoRA fine-tuned models built upon two mainstream architectures. We conduct an extensive evaluation of existing merging methods and further propose winner-take-all, a simple and hyperparameter-free approach that retains only the most dominant parameters across expert models. By preserving the critical parameters that govern model behavior and discarding weaker ones, our method avoids the information dilution inherent in averaging- or alignment-based strategies. Despite its simplicity, winner-take-all consistently outperforms existing approaches, offering both a new perspective on LoRA merging and a strong practical baseline for future research.

📄 PDF Abstract BibTeX arXiv:2607.15661

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuning

Similar Papers 제목 키워드 기반

HALLUCINOGEN: A Benchmark for Evaluating Object Hallucination in Large Visual-Language Models

2024-12-29 · Ashish Seth, Dinesh Manocha, Chirag Agarwal

Large Vision-Language Models (LVLMs) have demonstrated remarkable performance in performing complex multimodal tasks. However, they are still plagued by object hallucination: the misidentification or misclassification of…

HallucinationObjectObject HallucinationQuestion Answering+3

EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models

2025-09-24 · Botai Yuan, Yutian Zhou, Yingjie Wang, Fushuo Huo 외 arxiv

Recent benchmarks for medical Large Vision-Language Models (LVLMs) emphasize leaderboard accuracy, overlooking reliability and safety. We study sycophancy -- models' tendency to uncritically echo user-provided informatio…

Aloe-Vision: Robust Vision-Language Models for Healthcare

2026-06-25 · Jaume Guasch-Martí, Enrique Lopez-Cuena, Martín Suárez-Fernández, Jordi Bayarri-Planas 외 arxiv

Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications. However, progress is constrained by the…

Detecting and Evaluating Medical Hallucinations in Large Vision Language Models

2024-06-14 · Jiawei Chen, Dingkang Yang, Tong Wu, Yue Jiang 외

Large Vision Language Models (LVLMs) are increasingly integral to healthcare applications, including medical visual question answering and imaging report generation. While these models inherit the robust capabilities of …

HallucinationMedical Visual Question AnsweringQuestion AnsweringVisual Question Answering

Medical Large Vision Language Models with Multi-Image Visual Ability

2025-05-25 · Xikai Yang, Juzheng Miao, Yuchen Yuan, Jiaze Wang 외

Medical large vision-language models (LVLMs) have demonstrated promising performance across various single-image question answering (QA) benchmarks, yet their capability in processing multi-image clinical scenarios remai…

Question AnsweringVisual Question Answering (VQA)