paper-with-me

홈 › Papers

Evaluating and Understanding Model Editing for Medical Vision Language Models

2026-07-06 · Guli Zhu, Chenwei Wu, Liyue Shen arxiv

Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks and do not reflect realistic clinical domain requirements and variability. To address this, we introduce M3Bench, a clinically grounded benchmark for multimodal model editing that evaluates whether an edit remains reliable, precise, and generalizable under the challenges of image and text variation, modality and protocol shifts, clinical knowledge composition, and temporal progression. M3Bench contains 16,276 questions spanning diverse anatomy, modalities, and specialties, and supports both single and sequential edits. By evaluating 4 representative editors across 6 medical and general VLMs, we find that no method excels across all criteria. Gradient-based editors achieve strong transfer but suffer from catastrophic locality violations, whereas memory-based methods preserve locality but lack compositional generality and exhibit high backbone-dependent hyperparameter sensitivity. We further attribute these failures to the latent space geometry of VLMs and how different editing methods shift its landscape. Overall, M3Bench establishes a rigorous clinical stress test for multimodal model editing and offers actionable guidance for safer post-deployment adaptation. The benchmark is publicly available at https://github.com/BioMed-AI-Lab-U-Michgan/M3Bench .

📄 PDF Abstract BibTeX arXiv:2607.05310

Code (0)

등록된 구현이 없습니다.

Tasks

Clinical Knowledge

Similar Papers 제목 키워드 기반

MultiMedEdit: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQA

2025-08-09 · Shengtao Wen, Haodong Chen, Yadong Wang, Zhongying Pan 외 arxiv

Knowledge editing (KE) provides a scalable approach for updating factual knowledge in large language models without full retraining. While previous studies have demonstrated effectiveness in general domains and medical Q…

knowledge editingVisual Reasoning

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models

2025-08-07 · Dexuan Xu, Jieyi Wang, Zhongyan Chai, Yongzhi Cao 외 arxiv

Recent advances in multimodal large language models (MLLMs) have significantly improved medical AI, enabling it to unify the understanding of visual and textual information. However, as medical knowledge continues to evo…

Adversarial Robustnessknowledge editing

Multi-modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training

2021-05-24 · Jong Hak Moon, Hyungyung Lee, Woncheol Shin, Young-Hak Kim 외

Recently a number of studies demonstrated impressive performance on diverse vision-language multi-modal tasks such as image captioning and visual question answering by extending the BERT architecture with multi-modal pre…

Image CaptioningMedical Visual Question AnsweringMultimodal Deep LearningQuestion Answering+5

Can We Edit LLMs for Long-Tail Biomedical Knowledge?

2025-04-14 · Xinhao Yi, Jake Lever, Kevin Bryson, Zaiqiao Meng

Knowledge editing has emerged as an effective approach for updating large language models (LLMs) by modifying their internal knowledge. However, their application to the biomedical domain faces unique challenges due to t…

knowledge editing

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing

2026-06-23 · Qingxiao Li, Zikai Wang, Qingli Wang, Nan Xu arxiv

We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose image generation models, scientific image tasks require not only high-…

Image Super-ResolutionMultimodal ReasoningImage SegmentationImage Generation