paper-with-me

홈 › Papers

TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance

2025-09-30 · Yuyang Liu, Chuan Wen, Yihang Hu, Dinesh Jayaraman, Yang Gao arxiv

Designing dense rewards is crucial for reinforcement learning (RL), yet in robotics it often demands extensive manual effort and lacks scalability. One promising solution is to view task progress as a dense reward signal, as it quantifies the degree to which actions advance the system toward task completion over time. We present TimeRewarder, a simple yet effective reward learning method that derives progress estimation signals from passive videos, including robot demonstrations and human videos, by modeling temporal distances between frame pairs. We then demonstrate how TimeRewarder can supply step-wise proxy rewards to guide reinforcement learning. In our comprehensive experiments on ten challenging Meta-World tasks, we show that TimeRewarder dramatically improves RL for sparse-reward tasks, achieving nearly perfect success in 9/10 tasks with only 200,000 environment interactions per task. This approach outperformed previous methods and even the manually designed environment dense reward on both the final success rate and sample efficiency. Moreover, we show that TimeRewarder pretraining can exploit real-world human videos, highlighting its potential as a scalable approach to rich reward signals from diverse video sources.

📄 PDF Abstract BibTeX arXiv:2509.26627

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

On-Robot Reinforcement Learning with Goal-Contrastive Rewards

2024-10-25 · Ondrej Biza, Thomas Weng, Lingfeng Sun, Karl Schmeckpeper 외

Reinforcement Learning (RL) has the potential to enable robots to learn from their own actions in the real world. Unfortunately, RL can be prohibitively expensive, in terms of on-robot runtime, due to inefficient explora…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Reinforcement Learning from Passive Data via Latent Intentions

2023-04-10 · Dibya Ghosh, Chethan Bhateja, Sergey Levine

Passive observational data, such as human videos, is abundant and rich in information, yet remains largely untapped by current RL methods. Perhaps surprisingly, we show that passive data, despite not having reward or act…

reinforcement-learningReinforcement LearningValue prediction

Stage-Transition Dense Reward Modeling for Reinforcement Learning

2026-06-30 · Yang Yang, Bingjie Chen, Zihan Wang, Yizhe Li 외 arxiv

Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object conf…

Reinforcement Learning

VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

2022-09-30 · Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani 외

Reward and representation learning are two long-standing challenges for learning an expanding set of robot manipulation skills from sensory observations. Given the inherent cost and scarcity of in-domain, task-specific r…

Offline RLOpen-Ended Question AnsweringRepresentation LearningRobot Manipulation

Incentivizing Vision Language Models to Search for Long Video Question Answering

2026-07-03 · Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi 외 arxiv

We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek utilizes a natural language-driven sear…

Video Question AnsweringReinforcement Learning