paper-with-me

홈 › Papers

STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning

2025-09-28 · Yao Luan, Ni Mu, Yiqin Yang, Bo Xu, Qing-Shan Jia arxiv

Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning rewards directly from human preferences, enabling better alignment with human intentions. However, its effectiveness in multi-stage tasks, where agents sequentially perform sub-tasks (e.g., navigation, grasping), is limited by stage misalignment: Comparing segments from mismatched stages, such as movement versus manipulation, results in uninformative feedback, thus hindering policy learning. In this paper, we validate the stage misalignment issue through theoretical analysis and empirical experiments. To address this issue, we propose STage-AlIgned Reward learning (STAIR), which first learns a stage approximation based on temporal distance, then prioritizes comparisons within the same stage. Temporal distance is learned via contrastive learning, which groups temporally close states into coherent stages, without predefined task knowledge, and adapts dynamically to policy changes. Extensive experiments demonstrate STAIR's superiority in multi-stage tasks and competitive performance in single-stage tasks. Furthermore, human studies show that stages approximated by STAIR are consistent with human cognition, confirming its effectiveness in mitigating stage misalignment.

📄 PDF Abstract BibTeX arXiv:2509.23802

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningContrastive Learning

Similar Papers 제목 키워드 기반

Three-Stage Learning Unlocks Strong Performance in Simple Models for Long-Term Time Series Forecasting

2026-05-13 · Zhenan Yu, Guangxin Jiang, Jin Yang arxiv

Recent studies on long-term time series forecasting have shown that simple linear models and MLP-based predictors can achieve strong performance without increasingly complex architectures. However, many competitive basel…

Time Series Forecasting

StairMaster: Learning to Conquer Risky Hollow Stairs for Agile Quadrupedal Robots

2026-06-24 · Xincheng Tang, Youhan Xie, Zhengjie Shu, Wanyu Li 외 arxiv

Climbing hollow stairs remains a challenging problem for quadruped robots due to the high risk of leg trapping, severe depth sparsity, and high-frequency depth-sensing noise. In this paper, we propose StairMaster, a nove…

Reinforcement Learning

Training and Simulation of Quadrupedal Robot in Adaptive Stair Climbing for Indoor Firefighting: An End-to-End Reinforcement Learning Approach

2026-02-03 · Baixiao Huang, Baiyu Huang, Yu Hou arxiv

Quadruped robots are used for primary searches during the early stages of indoor fires. A typical primary search involves quickly and thoroughly looking for victims under hazardous conditions and monitoring flammable mat…

Reinforcement Learning

Learning Transferability: A Two-Stage Reinforcement Learning Approach for Enhancing Quadruped Robots' Performance in U-Shaped Stair Climbing

2026-02-16 · Baixiao Huang, Baiyu Huang, Yu Hou arxiv

Quadruped robots are employed in various scenarios in building construction. However, autonomous stair climbing across different indoor staircases remains a major challenge for robot dogs to complete building constructio…

Reinforcement Learning

Omnidirectional Humanoid Locomotion on Stairs via Unsafe Stepping Penalty and Sparse LiDAR Elevation Mapping

2026-03-09 · Yuzhi Jiang, Yujun Liang, Junhao Li, Han Ding 외 arxiv

Humanoid robots, characterized by numerous degrees of freedom and a high center of gravity, are inherently unstable. Safe omnidirectional locomotion on stairs requires both omnidirectional terrain perception and reliable…