paper-with-me

Papers

Anticipation-VLA: Solving Long-Horizon Embodied Tasks via Anticipation-based Subgoal Generation

2026-05-03 · Zhilong Zhang, Wenyu Luo, Haonan Wang, Yifei Sheng, Yidi Wang, Hanyuan Guo, Haoxiang Ren, Xinghao Du, Yuhan Che, Tongtong Cao, Lei Yuan, Yang Yu arxiv

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, enabling robots to perform tasks based on natural language instructions and current visual input. However, existing VLA models struggle with long-horizon tasks due to compounding errors. Prior methods decompose tasks into subtasks of fixed granularity, which cannot adapt to the varying complexity of execution states, limiting their robustness in long-horizon tasks. To overcome this, we introduce Anticipation Model, which adaptively and recursively generates future subgoals. This model continuously adapts as the task unfolds, adjusting future subgoals in response to evolving dynamics, facilitating more reliable planning paths. Building on this concept, we propose Anticipation-VLA, a hierarchical VLA model that leverages the anticipation model to generate actionable subgoals that guide VLA policy execution. We implement Anticipation-VLA with finetuning a Unified Multimodal Model (UMM) for high-level subgoal generation and a goal-conditioned VLA policy for low-level action execution. Experiments in both simulated and real-world robotic tasks demonstrate the effectiveness of Anticipation-VLA, highlighting the importance of adaptive and recursive subgoal generation for robust policy execution.

📄 PDF Abstract BibTeX arXiv:2605.01772

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reinforcement Learning with Anticipation: A Hierarchical Approach for Long-Horizon Tasks

2025-09-06 · Yang Yu arxiv

Solving long-horizon goal-conditioned tasks remains a significant challenge in reinforcement learning (RL). Hierarchical reinforcement learning (HRL) addresses this by decomposing tasks into more manageable sub-tasks, bu…

Hierarchical Reinforcement Learning

ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction

2025-11-26 · Qineng Wang, Wenlong Huang, Yu Zhou, Hang Yin 외 arxiv

Embodied cognition argues that intelligence arises from sensorimotor interaction rather than passive observation. It raises an intriguing question: do modern vision-language models (VLMs), trained largely in a disembodie…

Visual Question AnsweringAffordance Recognition

EVA: An Embodied World Model for Future Video Anticipation

2024-10-20 · Xiaowei Chi, Chun-Kai Fan, Hengyuan Zhang, Xingqun Qi 외

Video generation models have made significant progress in simulating future states, showcasing their potential as world simulators in embodied scenarios. However, existing models often lack robust understanding, limiting…

Language ModelingLanguage ModellingMixed RealityPrediction+3

Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning

2025-05-22 · Bosung Kim, Prithviraj Ammanabrolu

We introduce $\infty$-THOR, a new framework for long-horizon embodied tasks that advances long-context understanding in embodied AI. $\infty$-THOR provides: (1) a generation framework for synthesizing scalable, reproduci…

Long-Context Understanding

Long-horizon Embodied Planning with Implicit Logical Inference and Hallucination Mitigation

2024-09-24 · Siyuan Liu, Jiawei Du, Sicheng Xiang, Zibo Wang 외

Long-horizon embodied planning underpins embodied AI. To accomplish long-horizon tasks, one of the most feasible ways is to decompose abstract instructions into a sequence of actionable steps. Foundation models still fac…

DiversityHallucinationLanguage ModelingLanguage Modelling