paper-with-me

홈 › Papers

Reinforcing Action Policies by Prophesying

2025-11-25 · Jiahui Zhang, Ze Huang, Chun Gu, Zipei Ma, Li Zhang arxiv

Vision-Language-Action (VLA) policies excel in aligning language, perception, and robot control. However, most VLAs are trained purely by imitation, which overfits to demonstrations, and is brittle under distribution shift. Reinforcement learning (RL) directly optimizes task reward and thus addresses this misalignment, but real-robot interaction is expensive and conventional simulators are hard to engineer and transfer. We address both data efficiency and optimization stability in VLA post-training via a learned world model and an RL procedure tailored to flow-based action heads. Specifically, we first introduce Prophet, a unified action-to-video robot world model pretrained on large-scale, heterogeneous robot data to learn reusable action-outcome dynamics and then few-shot adapted to new robots, objects, and environments, yielding a rollout-ready simulator. Upon Prophet, we reinforce action policies with our proposed FlowScale, which couples Flow-GRPO with intrinsic stepwise reweighting to stabilize gradients. Together, our solution provides a practical, data- and compute-efficient path to VLA post-training. Experiments show 5-17% success gains on public benchmarks and 24-30% on real robots across diverse VLA backbones.

📄 PDF Abstract BibTeX arXiv:2511.20633

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Incentivized Bandit Learning with Self-Reinforcing User Preferences

2021-05-19 · Tianchen Zhou, Jia Liu, Chaosheng Dong, Jingyuan Deng

In this paper, we investigate a new multi-armed bandit (MAB) online learning model that considers real-world phenomena in many recommender systems: (i) the learning agent cannot pull the arms by itself and thus has to of…

Recommendation Systems

GSDrive: Reinforcing Driving Policies by Multi-mode Future Trajectory Probing with 3D Gaussian Splatting Environment

2026-04-30 · Ziang Guo, Chen Min, Xuefeng Zhang, Yixiao Zhou 외 arxiv

End-to-end (E2E) autonomous driving aims to directly map sensory observations to driving actions, but its real-world deployment is hindered by evolving data distributions and the high cost of continual annotation. While …

Reinforcement LearningAutonomous Driving

Fiscal Policy and Household Savings in Central Europe (Poland, Croatia, and Slovak Republic) -- A Markov Switching VAR with Covid Shock

2025-02-19 · Tuhin G M Al Mamun, Ehsanullah, Md Sharif Hassan, Mohd Faizal Yusof 외

This study investigates the effectiveness of fiscal policies on household consumption, disposable income, and the propensity to consume during the COVID-19 pandemic across Croatia, Slovakia, and Poland. The purpose is to…

PolicyBank: Evolving Policy Understanding for LLM Agents

2026-04-16 · Jihye Choi, Jinsung Yoon, Long T. Le, Somesh Jha 외 arxiv

LLM agents operating under organizational policies must comply with authorization constraints typically specified in natural language. In practice, such specifications inevitably contain ambiguities and logical or semant…

Reinforcing RCTs with Multiple Priors while Learning about External Validity

2021-12-16 · Frederico Finan, Demian Pouzo

This paper introduces a framework for incorporating prior information into the design of sequential experiments. These sources may include past experiments, expert opinions, or the experimenter's intuition. We model the …

valid