paper-with-me

Papers

HoRD: Robust Humanoid Control via History-Conditioned Reinforcement Learning and Online Distillation

2026-02-04 · Puyue Wang, Jiawei Hu, Yan Gao, Junyan Wang, Yu Zhang, Gillian Dobbie, Tao Gu, Wafa Johal, Ting Dang, Hong Jia arxiv

Humanoid robots can suffer significant performance drops under small changes in dynamics, task specifications, or environment setup. We propose HoRD, a two-stage learning framework for robust humanoid control under domain shift. First, we train a high-performance teacher policy via history-conditioned reinforcement learning, where the policy infers latent dynamics context from recent state--action trajectories to adapt online to diverse randomized dynamics. Second, we perform online distillation to transfer the teacher's robust control capabilities into a transformer-based student policy that operates on sparse root-relative 3D joint keypoint trajectories. By combining history-conditioned adaptation with online distillation, HoRD enables a single policy to adapt zero-shot to unseen domains without per-domain retraining. Extensive experiments show HoRD outperforms strong baselines in robustness and transfer, especially under unseen domains and external perturbations. Code and project page are available at https://tonywang-0517.github.io/hord/.

📄 PDF Abstract BibTeX arXiv:2602.04412

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Load-Aware Locomotion Control for Humanoid Robots in Industrial Transportation Tasks

2026-03-15 · Lequn Fu, Yijun Zhong, Xiao Li, Yibin Liu 외 arxiv

Humanoid robots deployed in industrial environments are required to perform load-carrying transportation tasks that tightly couple locomotion and manipulation. However, achieving stable and robust locomotion under varyin…

Reinforcement Learning

Chord-Conditioned Melody Harmonization with Controllable Harmonicity

2022-02-17 · Shangda Wu, Xiaobing Li, Maosong Sun

Melody harmonization has long been closely associated with chorales composed by Johann Sebastian Bach. Previous works rarely emphasised chorale generation conditioned on chord progressions, and there has been a lack of f…

Real-World Humanoid Locomotion with Reinforcement Learning

2023-03-06 · Ilija Radosavovic, Tete Xiao, Bike Zhang, Trevor Darrell 외

Humanoid robots that can autonomously operate in diverse environments have the potential to help address labour shortages in factories, assist elderly at homes, and colonize new planets. While classical controllers for h…

reinforcement-learningReinforcement Learning

PCHC: Enabling Preference Conditioned Humanoid Control via Multi-Objective Reinforcement Learning

2026-03-25 · Huanyu Li, Dewei Wang, Xinmiao Wang, Xinzhe Liu 외 arxiv

Humanoid robots often need to balance competing objectives, such as maximizing speed while minimizing energy consumption. While current reinforcement learning (RL) methods can master complex skills like fall recovery and…

Reinforcement Learning

Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control

2026-05-14 · Haozhe Jia, Honglei Jin, Yuan Zhang, Youcheng Fan 외 arxiv

Natural language is an intuitive interface for humanoid robots, yet streaming whole-body control requires control representations that are executable now and anticipatory of future physical transitions. Existing language…

Instruction Following