paper-with-me

홈 › Papers

Improving Robotic Manipulation Robustness via NICE Scene Surgery

2025-11-27 · Sajjad Pakdamansavoji, Mozhgan Pourkeshavarz, Adam Sigal, Zhiyuan Li, Rui Heng Yang, Amir Rasouli arxiv

Learning robust visuomotor policies for robotic manipulation remains a challenge in real-world settings, where visual distractors can significantly degrade performance and safety. In this work, we propose an effective and scalable framework, Naturalistic Inpainting for Context Enhancement (NICE). Our method minimizes out-of-distribution (OOD) gap in imitation learning by increasing visual diversity through construction of new experiences using existing demonstrations. By utilizing image generative frameworks and large language models, NICE performs three editing operations, object replacement, restyling, and removal of distracting (non-target) objects. These changes preserve spatial relationships without obstructing target objects and maintain action-label consistency. Unlike previous approaches, NICE requires no additional robot data collection, simulator access, or custom model training, making it readily applicable to existing robotic datasets. Using real-world scenes, we showcase the capability of our framework in producing photo-realistic scene enhancement. For downstream tasks, we use NICE data to finetune a vision-language model (VLM) for spatial affordance prediction and a vision-language-action (VLA) policy for object manipulation. Our evaluations show that NICE successfully minimizes OOD gaps, resulting in over 20% improvement in accuracy for affordance prediction in highly cluttered scenes. For manipulation tasks, success rate increases on average by 11% when testing in environments populated with distractors in different quantities. Furthermore, we show that our method improves visual robustness, lowering target confusion by 6%, and enhances safety by reducing collision rate by 7%.

📄 PDF Abstract BibTeX arXiv:2511.22777

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Real-time Monocular 2D and 3D Perception of Endoluminal Scenes for Controlling Flexible Robotic Endoscopic Instruments

2026-02-16 · Ruofeng Wei, Kai Chen, Yui Lun Ng, Yiyao Ma 외 arxiv

Endoluminal surgery offers a minimally invasive option for early-stage gastrointestinal and urinary tract cancers but is limited by surgical tools and a steep learning curve. Robotic systems, particularly continuum robot…

Learning and Reasoning with the Graph Structure Representation in Robotic Surgery

2020-07-07 · Mobarakol Islam, Lalithkumar Seenivasan, Lim Chwee Ming, Hongliang Ren

Learning to infer graph representations and performing spatial reasoning in a complex surgical environment can play a vital role in surgical scene understanding in robotic surgery. For this purpose, we develop an approac…

Edge ClassificationGraph GenerationScene Graph GenerationScene Segmentation+2

BASED: Bundle-Adjusting Surgical Endoscopic Dynamic Video Reconstruction using Neural Radiance Fields

2023-09-27 · Shreya Saha, Zekai Liang, Shan Lin, Jingpei Lu 외

Reconstruction of deformable scenes from endoscopic videos is important for many applications such as intraoperative navigation, surgical visual perception, and robotic surgery. It is a foundational requirement for reali…

NeRFVideo Reconstruction

Neural Rendering for Stereo 3D Reconstruction of Deformable Tissues in Robotic Surgery

2022-06-30 · Yuehao Wang, Yonghao Long, Siu Hin Fan, Qi Dou

Reconstruction of the soft tissues in robotic surgery from endoscopic stereo videos is important for many applications such as intra-operative navigation and image-guided robotic surgery automation. Previous works on thi…

3D ReconstructionNeural Rendering

Deep Reinforcement Learning Based Semi-Autonomous Control for Robotic Surgery

2022-04-11 · Ruiqi Zhu, Dandan Zhang, Benny Lo

In recent decades, the tremendous benefits surgical robots have brought to surgeons and patients have been witnessed. With the dexterous operation and the great precision, surgical robots can offer patients less recovery…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)