paper-with-me

Papers

AstraNav-World: World Model for Foresight Control and Consistency

2025-12-25 · Jintao Chen, Junjun Hu, Haochen Bai, Minghua Luo, Xinda Xue, Botao Ren, Chengyu Bai, Shichao Xie, Ziyi Chen, Fei Liu, Zedong Chu, Xiaolong Wu, Mu Xu, Shanghang Zhang arxiv

Embodied navigation in open, dynamic environments demands accurate foresight of how the world will evolve and how actions will unfold over time. We propose AstraNav-World, an end-to-end world model that jointly reasons about future visual states and action sequences within a unified probabilistic framework. Our framework integrates a diffusion-based video generator with a vision-language policy, enabling synchronized rollouts where predicted scenes and planned actions are updated simultaneously. Training optimizes two complementary objectives: generating action-conditioned multi-step visual predictions and deriving trajectories conditioned on those predicted visuals. This bidirectional constraint makes visual predictions executable and keeps decisions grounded in physically consistent, task-relevant futures, mitigating cumulative errors common in decoupled "envision-then-plan" pipelines. Experiments across diverse embodied navigation benchmarks show improved trajectory accuracy and higher success rates. Ablations confirm the necessity of tight vision-action coupling and unified training, with either branch removal degrading both prediction quality and policy reliability. In real-world testing, AstraNav-World demonstrated exceptional zero-shot capabilities, adapting to previously unseen scenarios without any real-world fine-tuning. These results suggest that AstraNav-World captures transferable spatial understanding and planning-relevant navigation dynamics, rather than merely overfitting to simulation-specific data distribution. Overall, by unifying foresight vision and control within a single generative model, we move closer to reliable, interpretable, and general-purpose embodied agents that operate robustly in open-ended real-world settings.

📄 PDF Abstract BibTeX arXiv:2512.21714

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Vision-Language-Action World Models for Autonomous Driving

2026-04-10 · Guoqing Wang, Pin Tang, Xiangxuan Ren, Guodongfang Zhao 외 arxiv

Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and control within a unified multimodal framework. However, they often la…

Reinforcement LearningAutonomous Driving

ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos

2026-01-08 · Rustin Soraki, Homanga Bharadhwaj, Ali Farhadi, Roozbeh Mottaghi arxiv

Humans can effortlessly anticipate how objects might move or change through interaction--imagining a cup being lifted, a knife slicing, or a lid being closed. We aim to endow computational systems with a similar ability …

3D Pose Estimation

Consistent World Models via Foresight Diffusion

2025-05-22 · Yu Zhang, Xingzhuo Guo, Haoran Xu, Mingsheng Long

Diffusion and flow-based models have enabled significant progress in generation tasks across various modalities and have recently found applications in world modeling. However, unlike typical generation tasks that encour…

AttributeDenoisingVideo Prediction

TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation

2026-06-09 · Yujie Zang, Yuhang Zheng, Xian Nie, Yupeng Zheng 외 arxiv

Contact-rich manipulation requires robots to continuously perceive and regulate evolving physical interactions under dynamic contact transitions or complex surface geometries. Recent imitation learning methods improve co…

Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach

2025-05-22 · Xiaoran Yin, Xu Luo, Hao Wu, Lianli Gao 외

The automatic control of mobile devices is essential for efficiently performing complex tasks that involve multiple sequential steps. However, these tasks pose significant challenges due to the limited environmental info…

Decision MakingNatural Language Understanding