paper-with-me

Papers

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation

2026-06-26 · Justin Yu, Andrew Goldberg, Kavish Kondap, Karim El-Refai, Ethan Ransing, Qianzhong Chen, Mac Schwager, Fred Shentu, Philipp Wu, Ken Goldberg arxiv

Scaling imitation learning requires large datasets, yet human teleoperation inevitably produces mixed-quality demonstrations containing hesitations and recoveries. Prior frame-level progress reward models supervise on absolute temporal progress proxies that suffer from label noise, or require costly human annotations to define subtask boundaries. We present WARP (Warp-Augmented Relative Progress), a novel fully self-supervised algorithm for learning dense, signed relative progress magnitudes directly from successful demonstrations. WARP generates per-frame progress targets via time-warp augmentations of demonstrations (variable playback speeds and reversals) and we train WARP-RM to predict the normalized elapsed time between input frames. Aggregating these predictions across overlapping windows yields a dense frame-level progress signal. We then introduce WARP-BC, which leverages these scalar reward estimates to upweight high-advantage action chunks during behavior cloning, where chunk-level advantage is obtained by aggregating per-frame rewards. We evaluate our approach on a physical bimanual robot system performing a long-horizon deformable object manipulation task: folding T-shirts from a random crumpled start. To evaluate policy robustness against suboptimal data, we construct training datasets of varying quality using episode length as a proxy for teleoperation sub-optimality. As the dataset is widened to admit more inefficiencies, WARP-BC maintains a 19/20 success rate compared to vanilla BC's collapse to 2/20, improving throughput by up to 18x. Furthermore, we evaluate a bottle-in-bin placement task in the real-world, as well as in a reproducible simulation of the task, where gains in success, speed, and throughput hold under paired significance tests, and we release all simulation code and evaluation artifacts. Project page: https://uynitsuj.github.io/warp-rm/

📄 PDF Abstract BibTeX arXiv:2606.28320

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WARP: On the Benefits of Weight Averaged Rewarded Policies

2024-06-24 · Alexandre Ramé, Johan Ferret, Nino Vieillard, Robert Dadashi 외

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) by encouraging their generations to have high rewards, using a reward model trained on human preferences. To prevent the forgetting of…

SynergyWarpNet: Attention-Guided Cooperative Warping for Neural Portrait Animation

2025-12-19 · Shihang Li, Zhiqiang Gong, Minming Ye, Yue Gao 외 arxiv

Recent advances in neural portrait animation have demonstrated remarked potential for applications in virtual avatars, telepresence, and digital content creation. However, traditional explicit warping approaches often st…

PG-VTON: A Novel Image-Based Virtual Try-On Method via Progressive Inference Paradigm

2023-04-18 · Naiyu Fang, Lemiao Qiu, Shuyou Zhang, Zili Wang 외

Virtual try-on is a promising computer vision topic with a high commercial value wherein a new garment is visually worn on a person with a photo-realistic effect. Previous studies conduct their shape and content inferenc…

Virtual Try-on

Autowarp: Learning a Warping Distance from Unlabeled Time Series Using Sequence Autoencoders

2018-10-23 · NeurIPS 2018 · Abubakar Abid, James Zou

Measuring similarities between unlabeled time series trajectories is an important problem in domains as diverse as medicine, astronomy, finance, and computer vision. It is often unclear what is the appropriate metric to …

AstronomyDynamic Time WarpingTime SeriesTime Series Analysis

Learning a Warping Distance from Unlabeled Time Series Using Sequence Autoencoders

2018-12-01 · NeurIPS 2018 12 · Abubakar Abid, James Y. Zou

Measuring similarities between unlabeled time series trajectories is an important problem in many domains such as medicine, economics, and vision. It is often unclear what is the appropriate metric to use because of the …

Dynamic Time WarpingTime SeriesTime Series Analysis