paper-with-me

홈 › Papers

Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery

2024-04-10 · Zohre Karimi, Shing-Hei Ho, Bao Thach, Alan Kuntz, Daniel S. Brown

Automating robotic surgery via learning from demonstration (LfD) techniques is extremely challenging. This is because surgical tasks often involve sequential decision-making processes with complex interactions of physical objects and have low tolerance for mistakes. Prior works assume that all demonstrations are fully observable and optimal, which might not be practical in the real world. This paper introduces a sample-efficient method that learns a robust reward function from a limited amount of ranked suboptimal demonstrations consisting of partial-view point cloud observations. The method then learns a policy by optimizing the learned reward function using reinforcement learning (RL). We show that using a learned reward function to obtain a policy is more robust than pure imitation learning. We apply our approach on a physical surgical electrocautery task and demonstrate that our method can perform well even when the provided demonstrations are suboptimal and the observations are high-dimensional point clouds. Code and videos available here: https://sites.google.com/view/lfdinelectrocautery

📄 PDF Abstract BibTeX arXiv:2404.07185

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingImitation LearningReinforcement Learning (RL)Sequential Decision Making

Similar Papers 제목 키워드 기반

Autonomous Vision-Guided Resection of Central Airway Obstruction

2025-02-25 · M. E. Smith, N. Yilmaz, T. Watts, P. M. Scheikl 외

Existing tracheal tumor resection methods often lack the precision required for effective airway clearance, and robotic advancements offer new potential for autonomous resection. We present a vision-guided, autonomous ap…

When Tracking Fails: Analyzing Failure Modes of SAM2 for Point-Based Tracking in Surgical Videos

2025-10-02 · Woowon Jang, Jiwon Im, Juseung Choi, Niki Rashidian 외 arxiv

Video object segmentation (VOS) models such as SAM2 offer promising zero-shot tracking capabilities for surgical videos using minimal user input. Among the available input types, point-based tracking offers an efficient …

Video Object Segmentation

D-Shape: Demonstration-Shaped Reinforcement Learning via Goal Conditioning

2022-10-26 · Caroline Wang, Garrett Warnell, Peter Stone

While combining imitation learning (IL) and reinforcement learning (RL) is a promising way to address poor sample efficiency in autonomous behavior acquisition, methods that do so typically assume that the requisite beha…

Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Seeing Through Smoke: Surgical Desmoking for Improved Visual Perception

2026-03-26 · Jingpei Lu, Fengyi Jiang, Xiaorui Zhang, Lingbo Jin 외 arxiv

Minimally invasive and robot-assisted surgery relies heavily on endoscopic imaging, yet surgical smoke produced by electrocautery and vessel-sealing instruments can severely degrade visual perception and hinder vision-ba…

Synthetic Data GenerationStereo Depth EstimationImage Reconstruction

Supervised Reward Inference

2025-02-25 · Will Schwarzer, Jordan Schneider, Philip S. Thomas, Scott Niekum

Existing approaches to reward inference from behavior typically assume that humans provide demonstrations according to specific models of behavior. However, humans often indicate their goals through a wide range of behav…