paper-with-me

Papers

Freeform Preference Learning for Robotic Manipulation

2026-06-30 · Marcel Torne, Anubha Mahajan, Abhijnya Bhat, Chelsea Finn arxiv

Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal. We introduce Freeform Preference Learning (FPL), a method for learning robot policies from freeform human preferences. Rather than asking annotators which of two trajectories is better overall, FPL lets them define natural-language preference axes, such as speed, safety, quality of placement, or carefulness, and provide pairwise preferences along each axis. These annotations are used to learn a language-conditioned reward model that maps a trajectory and preference label to an axis-specific reward. We use this model to train a reward-conditioned policy that optimizes across the multiple human-specified dimensions. Across four real-world and two simulated long-horizon manipulation tasks, FPL improves over sparse-reward and binary-preference methods by 38 percentage points. Beyond improved performance, FPL learns dense progress signals without explicit subtask segmentation, shows compositionality of behavior not present in the data, and allows users to steer the policy towards different behaviors at test time without retraining. Blog post with videos available at https://freeform-pl.github.io/fpl.website/

📄 PDF Abstract BibTeX arXiv:2606.32027

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive remanufacturing for freeform surface parts based on linear laser scanner and robotic laser cladding

2024-08-10 · Robotics and Computer-Integrated Manufacturing 2024 8 · Wei Ma, TianliangHu, Chengrui Zhang, Qizhi Chen

Freeform surface parts play a significant role in the aerospace industry, the mold- manufacturing industry and the automobile industry, and it is energy-saving, material-saving, time-saving and environmentally beneficial…

Preference Aligned Visuomotor Diffusion Policies for Deformable Object Manipulation

2026-02-10 · Marco Moletta, Michael C. Welle, Danica Kragic arxiv

Humans naturally develop preferences for how manipulation tasks should be performed, which are often subtle, personal, and difficult to articulate. Although it is important for robots to account for these preferences to …

Model-Free Co-Optimization of Manufacturable Sensor Layouts and Deformation Proprioception

2026-03-09 · Yingjun Tian, Guoxin Fang, Aoran Lyu, Xilong Wang 외 arxiv

Flexible sensors are increasingly employed in soft robotics and wearable devices to provide proprioception of freeform deformations.Although supervised learning can train shape predictors from sensor signals, prediction …

PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation

2023-06-06 · Runze Liu, Yali Du, Fengshuo Bai, Jiafei Lyu 외

In preference-based Reinforcement Learning (RL), obtaining a large number of preference labels are both time-consuming and costly. Furthermore, the queried human preferences cannot be utilized for the new tasks. In this …

Offline RLReinforcement Learning (RL)

Programmable Telescopic Soft Pneumatic Actuators for Deployable and Shape Morphing Soft Robots

2025-11-10 · Joel Kemp, Andre Farinha, David Howard, Krishna Manaswi Digumarti 외 arxiv

Soft Robotics presents a rich canvas for free-form and continuum devices capable of exerting forces in any direction and transforming between arbitrary configurations. However, there is no current way to tractably and di…