paper-with-me

홈 › Papers

MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation

2025-11-17 · Junjie Yang, Yuhao Yan, Gang Wu, Yuxuan Wang, Ruoyu Liang, Xinjie Jiang, Xiang Wan, Fenglei Fan, Yongquan Zhang, Feiwei Qin, Changmiao Wang arxiv

As Vision-Language Models (VLMs) increasingly gain traction in medical applications, clinicians are progressively expecting AI systems not only to generate textual diagnoses but also to produce corresponding medical images that integrate seamlessly into authentic clinical workflows. Despite the growing interest, existing medical visual benchmarks present notable limitations. They often rely on ambiguous queries that lack sufficient relevance to image content, oversimplify complex diagnostic reasoning into closed-ended shortcuts, and adopt a text-centric evaluation paradigm that overlooks the importance of image generation capabilities. To address these challenges, we introduce MedGEN-Bench, a comprehensive multimodal benchmark designed to advance medical AI research. MedGEN-Bench comprises 6,422 expert-validated image-text pairs spanning six imaging modalities, 16 clinical tasks, and 28 subtasks. It is structured into three distinct formats: Visual Question Answering, Image Editing, and Contextual Multimodal Generation. What sets MedGEN-Bench apart is its focus on contextually intertwined instructions that necessitate sophisticated cross-modal reasoning and open-ended generative outputs, moving beyond the constraints of multiple-choice formats. To evaluate the performance of existing systems, we employ a novel three-tier assessment framework that integrates pixel-level metrics, semantic text analysis, and expert-guided clinical relevance scoring. Using this framework, we systematically assess 10 compositional frameworks, 3 unified models, and 5 VLMs.

📄 PDF Abstract BibTeX arXiv:2511.13135

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answeringmultimodal generationImage GenerationImage Editing

Similar Papers 제목 키워드 기반

MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos

2025-07-08 · Rongsheng Wang, Junying Chen, Ke Ji, Zhenyang Cai 외 arxiv

Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical traini…

Video Generation

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

2026-05-09 · Weiren Zhao, Yi Dong, Cheng Chen arxiv

Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existing unified medical models typically treat understanding and generatio…

To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering

2024-03-04 · Giacomo Frisoni, Alessio Cocchieri, Alex Presepi, Gianluca Moro 외

Medical open-domain question answering demands substantial access to specialized knowledge. Recent efforts have sought to decouple knowledge from model parameters, counteracting architectural scaling and allowing for tra…

MedQAMMLUMultiple-choiceOpen-Domain Question Answering+1

MedGen3D: A Deep Generative Framework for Paired 3D Image and Mask Generation

2023-04-08 · Kun Han, Yifeng Xiong, Chenyu You, Pooya Khosravi 외

Acquiring and annotating sufficient labeled data is crucial in developing accurate and robust learning-based models, but obtaining such data can be challenging in many medical image segmentation tasks. One promising solu…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

Beyond Function-Level Search: Repository-Aware Dual-Encoder Code Retrieval with Adversarial Verification

2025-10-16 · Aofan Liu, Shiyuan Song, Haoxuan Li, Cehao Yang 외 arxiv

The escalating complexity of modern codebases has intensified the need for retrieval systems capable of interpreting cross-component change intents, a capability fundamentally absent in conventional function-level search…

Natural Language Queries