paper-with-me

홈 › Papers

Protective Policy Transfer

2020-12-11 · Wenhao Yu, C. Karen Liu, Greg Turk

Being able to transfer existing skills to new situations is a key capability when training robots to operate in unpredictable real-world environments. A successful transfer algorithm should not only minimize the number of samples that the robot needs to collect in the new environment, but also prevent the robot from damaging itself or the surrounding environment during the transfer process. In this work, we introduce a policy transfer algorithm for adapting robot motor skills to novel scenarios while minimizing serious failures. Our algorithm trains two control policies in the training environment: a task policy that is optimized to complete the task of interest, and a protective policy that is dedicated to keep the robot from unsafe events (e.g. falling to the ground). To decide which policy to use during execution, we learn a safety estimator model in the training environment that estimates a continuous safety level of the robot. When used with a set of thresholds, the safety estimator becomes a classifier for switching between the protective policy and the task policy. We evaluate our approach on four simulated robot locomotion problems and a 2D navigation problem and show that our method can achieve successful transfer to notably different environments while taking the robot's safety into consideration.

📄 PDF Abstract BibTeX arXiv:2012.06662

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Cross-Frequency Protective Emblem: Protective Options for Medical Units and Wounded Soldiers in the Context of (fully) Autonomous Warfare

2023-05-03 · Daniel C. Hinck, Jonas J. Schöttler, Maria Krantz, Katharina-Sophie Isleif 외

The protection of non-combatants in times of (fully) autonomous warfare raises the question of the timeliness of the international protective emblem. Incidents in the recent past indicate that it is becoming necessary to…

SafeFall: Learning Protective Control for Humanoid Robots

2025-11-23 · Ziyu Meng, Tengyu Liu, Le Ma, Yingying Wu 외 arxiv

Bipedal locomotion makes humanoid robots inherently prone to falls, causing catastrophic damage to the expensive sensors, actuators, and structural components of full-scale robots. To address this critical barrier to rea…

Reinforcement Learning

Beyond the Uncanny Valley: A Mixed-Method Investigation of Anthropomorphism in Protective Responses to Robot Abuse

2025-10-30 · Fan Yang, Lingyao Li, Yaxin Hu, Michael Rodgers 외 arxiv

Robots with anthropomorphic features are increasingly shaping how humans perceive and morally engage with them. Our research investigates how different levels of anthropomorphism influence protective responses to robot a…

Discovering Self-Protective Falling Policy for Humanoid Robot via Deep Reinforcement Learning

2025-12-01 · Diyuan Shi, Shangke Lyu, Donglin Wang arxiv

Humanoid robots have received significant research interests and advancements in recent years. Despite many successes, due to their morphology, dynamics and limitation of control policy, humanoid robots are prone to fall…

Reinforcement Learning

CoRPO: Adding a Correctness Bias to GRPO Improves Generalization

2025-11-06 · Anisha Garg, Claire Zhang, Nishit Neema, David Bick 외 arxiv

Group-Relative Policy Optimization (GRPO) has emerged as the standard for training reasoning capabilities in large language models through reinforcement learning. By estimating advantages using group-mean rewards rather …

Reinforcement Learning