paper-with-me

홈 › Papers

Language to Rewards for Robotic Skill Synthesis

2023-06-14 · Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montse Gonzalez Arenas, Hao-Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, Brian Ichter, Ted Xiao, Peng Xu, Andy Zeng, Tingnan Zhang, Nicolas Heess, Dorsa Sadigh, Jie Tan, Yuval Tassa, Fei Xia

Large language models (LLMs) have demonstrated exciting progress in acquiring diverse new capabilities through in-context learning, ranging from logical reasoning to code-writing. Robotics researchers have also explored using LLMs to advance the capabilities of robotic control. However, since low-level robot actions are hardware-dependent and underrepresented in LLM training corpora, existing efforts in applying LLMs to robotics have largely treated LLMs as semantic planners or relied on human-engineered control primitives to interface with the robot. On the other hand, reward functions are shown to be flexible representations that can be optimized for control policies to achieve diverse tasks, while their semantic richness makes them suitable to be specified by LLMs. In this work, we introduce a new paradigm that harnesses this realization by utilizing LLMs to define reward parameters that can be optimized and accomplish variety of robotic tasks. Using reward as the intermediate interface generated by LLMs, we can effectively bridge the gap between high-level language instructions or corrections to low-level robot actions. Meanwhile, combining this with a real-time optimizer, MuJoCo MPC, empowers an interactive behavior creation experience where users can immediately observe the results and provide feedback to the system. To systematically evaluate the performance of our proposed method, we designed a total of 17 tasks for a simulated quadruped robot and a dexterous manipulator robot. We demonstrate that our proposed method reliably tackles 90% of the designed tasks, while a baseline using primitive skills as the interface with Code-as-policies achieves 50% of the tasks. We further validated our method on a real robot arm where complex manipulation skills such as non-prehensile pushing emerge through our interactive system.

📄 PDF Abstract BibTeX arXiv:2306.08647

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningLogical ReasoningMuJoCo

Similar Papers 제목 키워드 기반

SLIM: Skill Learning with Multiple Critics

2024-02-01 · David Emukpere, Bingbing Wu, Julien Perez, Jean-Michel Renders

Self-supervised skill learning aims to acquire useful behaviors that leverage the underlying dynamics of the environment. Latent variable models, based on mutual information maximization, have been successful in this tas…

Hierarchical Reinforcement Learning

An Empowerment-based Solution to Robotic Manipulation Tasks with Sparse Rewards

2020-10-15 · Siyu Dai, Wei Xu, Andreas Hofmann, Brian Williams

In order to provide adaptive and user-friendly solutions to robotic manipulation, it is important that the agent can learn to accomplish tasks even if they are only provided with very sparse instruction signals. To addre…

Diversityreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Synthesis from Satisficing and Temporal Goals

2022-05-20 · Suguman Bansal, Lydia Kavraki, Moshe Y. Vardi, Andrew Wells

Reactive synthesis from high-level specifications that combine hard constraints expressed in Linear Temporal Logic LTL with soft constraints expressed by discounted-sum (DS) rewards has applications in planning and reinf…

Reinforcement Learning (RL)

Self-Supervised State-Control through Intrinsic Mutual Information Rewards

2019-09-25 · Rui Zhao, Volker Tresp, Wei Xu

Learning to discover useful skills without a manually-designed reward function would have many applications, yet is still a challenge for reinforcement learning. In this paper, we propose Mutual Information-based State-C…

OpenAI Gymreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Rewards, Experience Synthesis, and Continual Memory

2025-12-28 · Ken Huang, Jerry Huang arxiv

Reinforcement learning is increasingly used to transform large language models into agentic systems that act over long horizons, invoke tools, and manage memory under partial observability. While recent work has demonstr…

Reinforcement Learning