paper-with-me

Papers

PATCH: Action-Chunk-Conditioned Latent Patch Innovation Monitoring for Robot Manipulation

2026-06-15 · Yanan Zhou, Ranpeng Qiu, Yincong Chen, Jiajie Cui, Weiming Zhi arxiv

Learning-based manipulation policies have made substantial progress in real-world robot manipulation, particularly for short-horizon action generation. However, deployment in open workspaces remains fragile under unexpected local scene dynamics, such as moving objects, transient occlusions, or disturbances near the intended motion. Existing runtime monitors often rely on global observation anomalies, policy uncertainty, or frame-level visual changes, and struggle to distinguish task-relevant execution risk from benign visual variation. We introduce PATCH, an action-chunk-conditioned latent patch innovation monitor for deployment-time intervention. Given the active action chunk, PATCH defines a projected execution corridor, predicts latent patch evolution inside it, and accumulates persistent residuals unexplained by the robot's own motion. These residuals form a localized intervention signal that allows PATCH-Router to pause execution, select an available recovery source, and resume the original policy once localized innovation subsides. Experiments on real robot rollout data show that PATCH produces more stable and context-relevant triggers than competing runtime monitors. Real-robot deployment further demonstrates monitor-driven intervention and policy resumption for disturbance-aware manipulation. Project Page: https://yananzhou5555.github.io/PATCH/.

📄 PDF Abstract BibTeX arXiv:2606.16690

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

2026-07-29 · Yushan Liu, Peibo Sun, Xintao Chao, Zhenyang Yang 외 arxiv

Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore…

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse

2026-06-22 · Bole Ma, Jan Eitzinger, Harald Koestler, Gerhard Wellein arxiv

Multimodal agents repeatedly re-examine the same video frames, UI screenshots, and rendered artifacts as their context window slides and reasoning iterates, yet every look-back re-encodes from scratch, because prefix cac…

ConfAL-WM: Confidence-Guided Active Learning for Action-Conditioned World Models

2026-08-26 · Xiang Liu, Sen Cui, Changshui Zhang arxiv

Action-conditioned world models have become an important foundation for embodied prediction, planning, and synthetic data generation, but their errors under new task and scene distributions are often concentrated in loca…

Synthetic Data GenerationActive Learning

PatchBlock: A Lightweight Defense Against Adversarial Patches for Embedded EdgeAI Devices

2026-01-01 · Nandish Chattopadhyay, Abdul Basit, Amira Guesmi, Muhammad Abdullah Hanif 외 arxiv

Adversarial attacks pose a significant challenge to the reliable deployment of machine learning models in EdgeAI applications, such as autonomous driving and surveillance, which rely on resource-constrained devices for r…

Dimensionality ReductionAutonomous DrivingOutlier Detection

Shifted Chunk Transformer for Spatio-Temporal Representational Learning

2021-08-26 · NeurIPS 2021 12 · Xuefan Zha, Wentao Zhu, Tingxun Lv, Sen yang 외

Spatio-temporal representational learning has been widely adopted in various fields such as action recognition, video object segmentation, and action anticipation. Previous spatio-temporal representational learning appro…

Action AnticipationAction Recognitionimage-classificationImage Classification+3