paper-with-me

홈 › Papers

AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models

2026-04-20 · Tingzheng Jia, Kan Guo, Lanping Qian, Yongli Hu, Daxin Tian, Guixian Qu, Chunmian Lin, Baocai Yin, Jiapu Wang arxiv

Precision-critical manipulation requires both global trajectory organization and local execution correction, yet most vision-language-action (VLA) policies generate actions within a single unified space. This monolithic formulation forces macro-level transport and micro-level refinement to be optimized under the same objective, causing large motions to dominate learning while suppressing small but failure-critical corrective signals. In contrast, human manipulation is structured by global movement planning together with continuous local adjustment during execution. Motivated by this principle, we propose AnchorRefine, a hierarchical framework that factorizes VLA action modeling into trajectory anchor and residual refinement. The anchor planner predicts a coarse motion scaffold, while the refinement module corrects execution-level deviations to improve geometric and contact precision. We further introduce a decision-aware gripper refinement mechanism to better capture the discrete and boundary-sensitive nature of gripper control. Experiments on LIBERO, CALVIN, and real-robot tasks demonstrate that AnchorRefine consistently improves both regression-based and diffusion-based VLA backbones, yielding gains of up to 7.8% in simulation success rate and 18% in real-world success rate.

📄 PDF Abstract BibTeX arXiv:2604.17787

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency

2026-07-10 · Haohui Huang, Xi Yuan, Panpan Liao, Tao Teng 외 arxiv

Generated videos provide useful visual motion priors for robot manipulation, but their visual plausibility does not imply physical executability. A generated video usually lacks metric geometry, grasp grounding, robot ki…

Robot ManipulationVideo Generation

AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning

2026-07-03 · Qi Liu, Yabei Li, Hongsong Wang, Heng Zhang 외 arxiv

Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into executable continuous trajectories. Vision-Language-Action models have been introduc…

Trajectory PredictionAutonomous DrivingDecision Making

SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy

2026-04-20 · Wei Yao, Haohan Ma, Hongwen Zhang, Yunlian Sun 외 arxiv

Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied intelligence, due to severe data scarcity, complexities in multi-agent coordination, and limited generalization across …

AnchorRoute: Human Motion Synthesis with Interval-Routed Sparse Contro

2026-05-14 · Pengcheng Fang, Tengjiao Sun, Dongjie Fu, Xiaoyu Zhan 외 arxiv

Sparse anchors provide a compact interface for human motion authoring: users specify a few root positions, planar trajectory samples, or body-point targets, while the system synthesizes the full-body motion that complete…

Motion Synthesis

AnchorVLA: Anchored Diffusion for Efficient End-to-End Mobile Manipulation

2026-04-02 · Jia Syuen Lim, Zhizhen Zhang, Peter Bohm, Brendan Tidd 외 arxiv

A central challenge in mobile manipulation is preserving multiple plausible action models while remaining reactive during execution. A bottle in a cluttered scene can often be approached and grasped in multiple valid way…