paper-with-me

홈 › Papers

EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities

2025-10-31 · Travis Davies, Yiqi Huang, Alexi Gladstone, Yunxin Liu, Xiang Chen, Heng Ji, Huxian Liu, Luhui Hu arxiv

Implicit policies parameterized by generative models, such as Diffusion Policy, have become the standard for policy learning and Vision-Language-Action (VLA) models in robotics. However, these approaches often suffer from high computational cost, exposure bias, and unstable inference dynamics, which lead to divergence under distribution shifts. Energy-Based Models (EBMs) address these issues by learning energy landscapes end-to-end and modeling equilibrium dynamics, offering improved robustness and reduced exposure bias. Yet, policies parameterized by EBMs have historically struggled to scale effectively. Recent work on Energy-Based Transformers (EBTs) demonstrates the scalability of EBMs to high-dimensional spaces, but their potential for solving core challenges in physically embodied models remains underexplored. We introduce a new energy-based architecture, EBT-Policy, that solves core issues in robotic and real-world settings. Across simulated and real-world tasks, EBT-Policy consistently outperforms diffusion-based policies, while requiring less training and inference computation. Remarkably, on some tasks it converges within just two inference steps, a 50x reduction compared to Diffusion Policy's 100. Moreover, EBT-Policy exhibits emergent capabilities not seen in prior models, such as zero-shot recovery from failed action sequences using only behavior cloning and without explicit retry training. By leveraging its scalar energy for uncertainty-aware inference and dynamic compute allocation, EBT-Policy offers a promising path toward robust, generalizable robot behavior under distribution shifts.

📄 PDF Abstract BibTeX arXiv:2510.27545

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Neuromorphic Reinforcement Learning Framework for Efficient Pathfinding in Robotic Mobile Fulfillment Systems

2026-06-18 · Junzhe Xu, Zecui Zeng, Lusong Li, Yuetong Fang 외 arxiv

Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robotic Mobile Fulfillment Systems (RMFS) a challenging problem for conventional search- and rule-based methods,…

Knowledge DistillationReinforcement Learning

Trading off regional and overall energy system design flexibility in the net-zero transition

2023-12-18 · Koen van Greevenbroek, Aleksander Grochowicz, Marianne Zeyringer, Fred Espen Benth

The transition to net-zero emissions in Europe is determined by a patchwork of country-level and EU-wide policy, creating coordination challenges in an interconnected system. We use an optimisation model to map out near-…

EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning

2025-02-18 · Xiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu 외

Large Language Models (LLMs) have shown impressive reasoning capabilities in well-defined problems with clear solutions, such as mathematics and coding. However, they still struggle with complex real-world scenarios like…

NavigateReinforcement Learning (RL)

Quasi-Equivalence Discovery for Zero-Shot Emergent Communication

2021-03-14 · Kalesha Bullard, Douwe Kiela, Franziska Meier, Joelle Pineau 외

Effective communication is an important skill for enabling information exchange in multi-agent settings and emergent communication is now a vibrant field of research, with common settings involving discrete cheap-talk ch…

PH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative Dynamics

2026-05-18 · Xueyu Luan, Chenwei Shi arxiv

World models built on recurrent state space architectures enable efficient latent imagination, yet remain physically unstructured, producing dynamics that violate conservation and dissipative principles. We introduce a u…