paper-with-me

홈 › Papers

HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation

2025-06-26 · Diego Biagini, Nassir Navab, Azade Farshad

Surgical Video Synthesis has emerged as a promising research direction following the success of diffusion models in general-domain video generation. Although existing approaches achieve high-quality video generation, most are unconditional and fail to maintain consistency with surgical actions and phases, lacking the surgical understanding and fine-grained guidance necessary for factual simulation. We address these challenges by proposing HieraSurg, a hierarchy-aware surgical video generation framework consisting of two specialized diffusion models. Given a surgical phase and an initial frame, HieraSurg first predicts future coarse-grained semantic changes through a segmentation prediction model. The final video is then generated by a second-stage model that augments these temporal segmentation maps with fine-grained visual features, leading to effective texture rendering and integration of semantic information in the video space. Our approach leverages surgical information at multiple levels of abstraction, including surgical phase, action triplets, and panoptic segmentation maps. The experimental results on Cholecystectomy Surgical Video Generation demonstrate that the model significantly outperforms prior work both quantitatively and qualitatively, showing strong generalization capabilities and the ability to generate higher frame-rate videos. The model exhibits particularly fine-grained adherence when provided with existing segmentation maps, suggesting its potential for practical surgical applications.

📄 PDF Abstract BibTeX arXiv:2506.21287

Code (0)

등록된 구현이 없습니다.

Tasks

Panoptic SegmentationSegmentationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy

2026-03-31 · Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma 외 arxiv

Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes. Computer-assisted systems such as surgical visual question answering (VQA…

Visual Question AnsweringVideo Question Answering

SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary

2026-05-20 · Jingyi He, Yue Zhou, Long Bai, Kun Yuan 외 arxiv

Understanding surgical workflow in real time is fundamental for intelligent surgical embodiment, where AI systems continuously perceive and respond as surgery proceeds. In the operating room, critical decisions depend on…

SurGen: Text-Guided Diffusion Model for Surgical Video Generation

2024-08-26 · Joseph Cho, Samuel Schmidgall, Cyril Zakka, Mrudang Mathur 외

Diffusion-based video generation models have made significant strides, producing outputs with improved visual fidelity, temporal coherence, and user control. These advancements hold great promise for improving surgical e…

Video Generation

Rethinking Surgical Smoke: A Smoke-Type-Aware Laparoscopic Video Desmoking Method and Dataset

2025-12-02 · Qifan Liang, Junlin Li, Zhen Han, Xihao Wang 외 arxiv

Electrocautery or lasers will inevitably generate surgical smoke, which hinders the visual guidance of laparoscopic videos for surgical procedures. The surgical smoke can be classified into different types based on its m…

Video Reconstruction

SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation

2024-12-18 · Tong Chen, Shuya Yang, Junyi Wang, Long Bai 외

Surgical video generation can enhance medical education and research, but existing methods lack fine-grained motion control and realism. We introduce SurgSora, a framework that generates high-fidelity, motion-controllabl…

Optical Flow EstimationVideo Generation