paper-with-me

홈 › Papers

HORIZON: Recoverability-Governed Curriculum for Physical-Domain Scaling

2026-06-03 · Chenhao Bai, Liqin Lu, Kaijun Wang, Hui Chen, Jin-Chuan Shi, Yuyang Liu, Hao Chen, Chunhua Shen arxiv

Scaling robust robot policies requires more than broader randomization, because physical-domain experience must remain organized and learnable throughout training. We study when a policy can benefit from harder physics and identify recoverability as a central constraint in on-policy physical-domain scaling. In on-policy training, new dynamics are useful only insofar as they remain close enough to the current policy to generate corrective on-policy data, rather than collapsing rollouts into unrecoverable failures. Using quadruped locomotion as a physically demanding benchmark for embodied generalization, we introduce HORIZON, a checkpointed frontier curriculum that expands physical domains only within the current policy's recoverable boundary. HORIZON uses rollback and boundary refinement to govern each expansion step, turning fixed randomization into a continual process of physical-domain growth. Experiments reveal three regularities of physical-domain expansion. First, direct domain widening is uneven across physical axes and often unlearnable without staged ordering. Second, domain composition is non-monotonic, and adding more domains beyond a compact core can dilute recoverable joint samples and reduce overall robustness. Third, offline distillation of isolated experts cannot substitute for the joint interaction generated by on-policy curriculum. Together, these results frame physical-domain generalization as a continual growth problem for embodied control, with recoverability as the organizing principle for on-policy expansion.

📄 PDF Abstract BibTeX arXiv:2606.05143

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Generalization

Similar Papers 제목 키워드 기반

Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents

2026-01-29 · Sri Vatsa Vuddanti, Satwik Kumar Chittiprolu arxiv

Language model agents often appear capable of self-recovery after failing tool call executions, yet this behavior lacks a formal explanation. We present a predictive theory that resolves this gap by showing that recovera…

Learning Neural PDE Solvers with Parameter-Guided Channel Attention

2023-04-27 · Makoto Takamoto, Francesco Alesiani, Mathias Niepert

Scientific Machine Learning (SciML) is concerned with the development of learned emulators of physical systems governed by partial differential equations (PDE). In application domains such as weather forecasting, molecul…

PDE Surrogate ModelingWeather Forecasting

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

2026-02-16 · Yu Huang, Zixin Wen, Yuejie Chi, Yuting Wei 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it remains a mystery how rewards based solely on final outcomes can help overcome the lo…

Reinforcement Learning

Learning Tennis Strategy Through Curriculum-Based Dueling Double Deep Q-Networks

2025-12-20 · Vishnu Mohan arxiv

Tennis strategy optimization is a challenging sequential decision-making problem involving hierarchical scoring, stochastic outcomes, long-horizon credit assignment, physical fatigue, and adaptation to opponent skill. I …

Reinforcement Learning

h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning

2025-10-08 · Sumeet Ramesh Motwani, Alesia Ivanova, Ziyang Cai, Philip Torr 외 arxiv

Large language models excel at short-horizon reasoning tasks, but performance drops as reasoning horizon lengths increase. Existing approaches to combat this rely on inference-time scaffolding or costly step-level superv…

Reinforcement Learning