paper-with-me

홈 › Papers

SurgCUT3R: Surgical Scene-Aware Continuous Understanding of Temporal 3D Representation

2026-03-07 · Kaiyuan Xu, Fangzhou Hong, Daniel Elson, Baoru Huang arxiv

Reconstructing surgical scenes from monocular endoscopic video is critical for advancing robotic-assisted surgery. However, the application of state-of-the-art general-purpose reconstruction models is constrained by two key challenges: the lack of supervised training data and performance degradation over long video sequences. To overcome these limitations, we propose SurgCUT3R, a systematic framework that adapts unified 3D reconstruction models to the surgical domain. Our contributions are threefold. First, we develop a data generation pipeline that exploits public stereo surgical datasets to produce large-scale, metric-scale pseudo-ground-truth depth maps, effectively bridging the data gap. Second, we propose a hybrid supervision strategy that couples our pseudo-ground-truth with geometric self-correction to enhance robustness against inherent data imperfections. Third, we introduce a hierarchical inference framework that employs two specialized models to effectively mitigate accumulated pose drift over long surgical videos: one for global stability and one for local accuracy. Experiments on the SCARED and StereoMIS datasets demonstrate that our method achieves a competitive balance between accuracy and efficiency, delivering near state-of-the-art but substantially faster pose estimation and offering a practical and effective solution for robust reconstruction in surgical environments. Project page: https://chumo-xu.github.io/SurgCUT3R-ICRA26/.

📄 PDF Abstract BibTeX arXiv:2603.06971

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionPose Estimation

Similar Papers 제목 키워드 기반

Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes

2025-10-28 · Jonas Hein, Lazaros Vlachopoulos, Maurits Geert Laurent Olthof, Bastian Sigrist 외 arxiv

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual…

Sound Source LocalizationScene UnderstandingPoint Clouds

Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry

2024-11-17 · Wenjun Hou, Yi Cheng, Kaishuai Xu, Yan Hu 외

Comprehensively understanding surgical scenes in Surgical Visual Question Answering (Surgical VQA) requires reasoning over multiple objects. Previous approaches address this task using cross-modal fusion strategies to en…

Question AnsweringScene UnderstandingVisual Question AnsweringVisual Question Answering (VQA)

Toward Real-Time Surgical Scene Segmentation via a Spike-Driven Video Transformer with Spike-Informed Pretraining

2025-12-24 · Shihao Zou, Jingjing Li, Wei Ji, Jincai Huang 외 arxiv

Modern surgical systems increasingly rely on intelligent scene understanding to improve intra-operative safety and situational awareness, with surgical scene segmentation playing a fundamental role in fine-grained surgic…

Knowledge DistillationScene UnderstandingScene Segmentation

SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting

2025-06-29 · Yiming Huang, Long Bai, Beilei Cui, Kun Yuan 외

In contemporary surgical research and practice, accurately comprehending 3D surgical scenes with text-promptable capabilities is particularly crucial for surgical planning and real-time intra-operative guidance, where pr…

3D ReconstructionScene Understanding

SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy

2026-07-05 · Lei Song, Yonghao Long, Mengya Xu, Jiayi Geng 외 arxiv

Surgical automation is being increasingly studied, yet bridging visual scene understanding with autonomous action planning remains a fundamental challenge. While much research effort has been made on scene perception (e.…

Scene UnderstandingScene Segmentation