paper-with-me

홈 › Papers

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

2026-08-19 · Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, Jian Wu, Xian Wu, Zuozhu Liu arxiv

Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absence of comprehensive training and evaluation benchmarks, and the lack of broadly validated unified medical model. To address these gaps, we present a comprehensive foundation for medical UAG. First, we construct MedUAGCorpus, the largest unified medical understanding and generation dataset to date, comprising over 6 million instances across 14 imaging modalities. Second, we introduce MedUAGBench, a systematic benchmark that expands medical generation evaluation to 12 diverse tasks under standardized protocols. Finally, leveraging these resources, we develop MedUAG, an end-to-end trained unified medical model. Extensive experiments demonstrate that MedUAG achieves strong performance across a wide array of understanding and generation tasks, establishing a competitive baseline and paving the way for next-generation medical multimodal systems.

📄 PDF Abstract BibTeX arXiv:2608.18937

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

2025-10-17 · Junzhi Ning, Wei Li, Cheng Tang, Jiashi Lin 외 arxiv

Medical workflows routinely combine reading images with producing visual and textual outputs, making both image understanding and generation central to medical AI. Most existing systems, however, address these abilities …

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

2026-05-09 · Weiren Zhao, Yi Dong, Cheng Chen arxiv

Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existing unified medical models typically treat understanding and generatio…

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing

2026-06-23 · Qingxiao Li, Zikai Wang, Qingli Wang, Nan Xu arxiv

We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose image generation models, scientific image tasks require not only high-…

Image Super-ResolutionMultimodal ReasoningImage SegmentationImage Generation

MedXChat: A Unified Multimodal Large Language Model Framework towards CXRs Understanding and Generation

2023-12-04 · Ling Yang, Zhanyu Wang, Zhenghao Chen, Xinyu Liang 외

Multimodal Large Language Models (MLLMs) have shown success in various general image processing tasks, yet their application in medical imaging is nascent, lacking tailored models. This study investigates the potential o…

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model+4

Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

2025-06-08 · LASA Team, Weiwen Xu, Hou Pong Chan, Long Li 외

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in understanding common visual elements, largely due to their large-scale datasets and advanced training strategies. However, their effec…

Medical Report GenerationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)