paper-with-me

홈 › Papers

Visuomotor Grasping with World Models for Surgical Robots

2025-08-15 · Hongbin Lin, Bin Li, Kwok Wai Samuel Au arxiv

Grasping is a fundamental task in robot-assisted surgery (RAS), and automating it can reduce surgeon workload while enhancing efficiency, safety, and consistency beyond teleoperated systems. Most prior approaches rely on explicit object pose tracking or handcrafted visual features, limiting their generalization to novel objects, robustness to visual disturbances, and the ability to handle deformable objects. Visuomotor learning offers a promising alternative, but deploying it in RAS presents unique challenges, such as low signal-to-noise ratio in visual observations, demands for high safety and millimeter-level precision, as well as the complex surgical environment. This paper addresses three key challenges: (i) sim-to-real transfer of visuomotor policies to ex vivo surgical scenes, (ii) visuomotor learning using only a single stereo camera pair -- the standard RAS setup, and (iii) object-agnostic grasping with a single policy that generalizes to diverse, unseen surgical objects without retraining or task-specific models. We introduce Grasp Anything for Surgery V2 (GASv2), a visuomotor learning framework for surgical grasping. GASv2 leverages a world-model-based architecture and a surgical perception pipeline for visual observations, combined with a hybrid control system for safe execution. We train the policy in simulation using domain randomization for sim-to-real transfer and deploy it on a real robot in both phantom-based and ex vivo surgical settings, using only a single pair of endoscopic cameras. Extensive experiments show our policy achieves a 65% success rate in both settings, generalizes to unseen objects and grippers, and adapts to diverse disturbances, demonstrating strong performance, generality, and robustness.

📄 PDF Abstract BibTeX arXiv:2508.11200

Code (0)

등록된 구현이 없습니다.

Tasks

Pose Tracking

Similar Papers 제목 키워드 기반

World Models for General Surgical Grasping

2024-05-28 · Hongbin Lin, Bin Li, Chun Wai Wong, Juan Rojas 외

Intelligent vision control systems for surgical robots should adapt to unknown and diverse objects while being robust to system disturbances. Previous methods did not meet these requirements due to mainly relying on pose…

Deep Reinforcement LearningPose Estimation

Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations

2026-04-08 · Chao Tang, Jiacheng Xu, Haofei Lu, Bolin Zou 외 arxiv

Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vast diversity of objects and tasks. Existing methods are either constr…

Video Generation

Learning a visuomotor controller for real world robotic grasping using simulated depth images

2017-06-14 · Ulrich Viereck, Andreas ten Pas, Kate Saenko, Robert Platt

We want to build robots that are useful in unstructured real world applications, such as doing work in the household. Grasping in particular is an important skill in this domain, yet it remains a challenge. One of the ke…

Robotic Grasping

Enhancing a Neurocognitive Shared Visuomotor Model for Object Identification, Localization, and Grasping With Learning From Auxiliary Tasks

2020-09-26 · Matthias Kerzel, Fares Abawi, Manfred Eppe, Stefan Wermter

We present a follow-up study on our unified visuomotor neural model for the robotic tasks of identifying, localizing, and grasping a target object in a scene with multiple objects. Our Retinanet-based model enables end-t…

Learning Surgical Robotic Manipulation with 3D Spatial Priors

2026-03-04 · Yu Sheng, Lidian Wang, Xiaomeng Chu, Jiajun Deng 외 arxiv

Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or…