paper-with-me

홈 › Papers

Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents

2026-06-22 · Haoran Zhang, Yifu Lu, Boyang Wang, Xuhui Kang, Yen-Ling Kuo, Zezhou Cheng, Mengdi Wang, Odest Chadwicke Jenkins arxiv

Long-horizon tasks are common in real-world robotic deployments, yet failure detection for such tasks remains underexplored. Detecting failures in long-horizon robotic tasks is particularly challenging because failure onset is often ambiguous and dense temporal annotations are typically unavailable. We present Foresight, a failure detection framework that monitors manipulation trajectories using latent representations from an action-conditioned world model. Foresight is trained using only final task-level success or failure labels. By leveraging predictive world-model embeddings, our method provides a unified framework for failure detection across different policies. We further use functional conformal prediction (FCP) to calibrate detection thresholds adaptively. We evaluate Foresight with state-of-the-art vision-language-action policies in simulation on LIBERO-Long, ManiSkill-Long, and BEHAVIOR-1K, compare it against state-of-the-artfailure detection methods, and validate it on real robots with three long-horizon tasks on a ReactorX-200 arm and one task on a Franka arm. Our results suggest that action-conditioned world-model embeddings provide a scalable representation for reliable failure monitoring in long-horizon manipulation.

📄 PDF Abstract BibTeX arXiv:2606.23085

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems

2026-05-09 · Boxuan Zhang, Jianing Zhu, Zeru Shi, Dongfang Liu 외 arxiv

LLM-based multi-agent systems are increasingly deployed on long-horizon tasks, but a single decisive error is often accepted by downstream agents and cascades into trajectory-level failure. Existing work frames this as \…

Reinforcement Learning

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

2025-12-10 · Minghui Lin, Pengxiang Ding, Shu Wang, Zifeng Zhuang 외 arxiv

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation …

Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation

2026-03-13 · Minghao Jin, Mozheng Liao, Mingfei Han, Zhihui Li 외 arxiv

Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introduces visual redundancy and accumulates er…

UF-RNN: Real-Time Adaptive Motion Generation Using Uncertainty-Driven Foresight Prediction

2025-10-11 · Hyogo Hiruma, Hiroshi Ito, Tetsuya Ogata arxiv

Training robots to operate effectively in environments with uncertain states, such as ambiguous object properties or unpredictable interactions, remains a longstanding challenge in robotics. Imitation learning methods ty…

STORM: Search-Guided Generative World Models for Robotic Manipulation

2025-12-20 · Wenjun Lin, Jensen Zhang, Kaitong Cai, Keze Wang arxiv

We present STORM (Search-Guided Generative World Models), a novel framework for spatio-temporal reasoning in robotic manipulation that unifies diffusion-based action generation, conditional video prediction, and search-b…

Video Prediction