paper-with-me

홈 › Papers

MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos

2025-07-08 · Rongsheng Wang, Junying Chen, Ke Ji, Zhenyang Cai, Shunian Chen, Yunjin Yang, Benyou Wang arxiv

Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical training, education, and simulation, requiring not only high visual fidelity but also strict medical accuracy. However, current models often produce unrealistic or erroneous content when applied to medical prompts, largely due to the lack of large-scale, high-quality datasets tailored to the medical domain. To address this gap, we introduce MedVideoCap-55K, the first large-scale, diverse, and caption-rich dataset for medical video generation. It comprises over 55,000 curated clips spanning real-world medical scenarios, providing a strong foundation for training generalist medical video generation models. Built upon this dataset, we develop MedGen, which achieves leading performance among open-source models and rivals commercial systems across multiple benchmarks in both visual quality and medical accuracy. We hope our dataset and model can serve as a valuable resource and help catalyze further research in medical video generation. Our code and data is available at https://github.com/FreedomIntelligence/MedGen

📄 PDF Abstract BibTeX arXiv:2507.05675

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

2026-05-09 · Weiren Zhao, Yi Dong, Cheng Chen arxiv

Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existing unified medical models typically treat understanding and generatio…

MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation

2025-11-17 · Junjie Yang, Yuhao Yan, Gang Wu, Yuxuan Wang 외 arxiv

As Vision-Language Models (VLMs) increasingly gain traction in medical applications, clinicians are progressively expecting AI systems not only to generate textual diagnoses but also to produce corresponding medical imag…

Visual Question Answeringmultimodal generationImage GenerationImage Editing

MedGen3D: A Deep Generative Framework for Paired 3D Image and Mask Generation

2023-04-08 · Kun Han, Yifeng Xiong, Chenyu You, Pooya Khosravi 외

Acquiring and annotating sufficient labeled data is crucial in developing accurate and robust learning-based models, but obtaining such data can be challenging in many medical image segmentation tasks. One promising solu…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering

2024-03-04 · Giacomo Frisoni, Alessio Cocchieri, Alex Presepi, Gianluca Moro 외

Medical open-domain question answering demands substantial access to specialized knowledge. Recent efforts have sought to decouple knowledge from model parameters, counteracting architectural scaling and allowing for tra…

MedQAMMLUMultiple-choiceOpen-Domain Question Answering+1

Ascle: A Python Natural Language Processing Toolkit for Medical Text Generation

2023-11-28 · Rui Yang, Qingcheng Zeng, Keen You, Yujie Qiao 외

This study introduces Ascle, a pioneering natural language processing (NLP) toolkit designed for medical text generation. Ascle is tailored for biomedical researchers and healthcare professionals with an easy-to-use, all…

Machine TranslationQuestion AnsweringSentence segmentationText Generation+3