paper-with-me

홈 › Papers

Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models

2026-07-29 · Mingyang Sun, Jiude Wei, Xiujian Liang, Qichen He, Donglin Wang, Cewu Lu, Jianhua Sun arxiv

Current Vision-Language-Action (VLA) models rely mainly on 2D inputs, neglecting the rich object structural information and commonsense knowledge inherent in the 3D physical world. This deficiency restricts their spatial awareness and adaptability for complex, high-precision manipulation. To bridge this crucial gap, we construct a Concept Expert module for VLA to build executable Analytic Concepts that represent objects as explicit, programmatic blueprints. Our mechanism operates in two synergistic phases: First, prior to VLA inference, the Concept Expert leverages 3D information from Vision Foundation Models (VFMs) to estimate the initial kinematic and structural parameters. Second, throughout the manipulation process, the VLA model utilizes its inherent capability to dynamically track the dynamic concept parameters, continuously aligning them with observational changes to ensure persistent accuracy. Once established, the Analytic Concepts provide explicit, high-quality guidance for VLA fine-tuning through (1) dense, programmatic manipulation rewards and (2) precise spatial guidance. This formulation allows VLA models to learn physically grounded interaction behaviors while maintaining end-to-end learning flexibility. Our experimental results show consistent improvements in success rate and learning efficiency across supervised and reinforcement learning settings, demonstrating the effectiveness of structured, concept-based guidance for VLA post-training.

📄 PDF Abstract BibTeX arXiv:2607.26513

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

IOI: Decoupling Kinematics and Physics for Interactive World Models

2026-06-22 · Chengyu Bai, Peidong Jia, Tiecheng Guo, Yukai Wang 외 arxiv

Developing generalist embodied agents requires interactive environments providing visually realistic feedback and accurate action-conditioned dynamics. Interactive world models address this by simulating such complex dyn…

Zero-shot GeneralizationVideo Generation

Minimum Radiative Heat and Propellant Aerocapture Guidance with Attitude Kinematics Constraints

2024-11-05 · Enrico Marco Zucchelli, Erwin Mooij

Aerocapture leverages atmospheric drag to convert a spacecraft's hyperbolic trajectory into a bound orbit. For some aerocapture missions, heating due to the radiation of high temperature gases in the shock-layer can be m…

Extract, Denoise and Enforce: Evaluating and Improving Concept Preservation for Text-to-Text Generation

2021-04-18 · EMNLP 2021 11 · Yuning Mao, Wenchang Ma, Deren Lei, Jiawei Han 외

Prior studies on text-to-text generation typically assume that the model could figure out what to attend to in the input and what to include in the output via seq2seq learning, with only the parallel training data and no…

Conditional Text GenerationDenoisingText Generation

KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition

2026-03-18 · Gaoge Han, Zhengqing Gao, Ziwen Li, Jiaxin Huang 외 arxiv

In this paper, we introduce a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes (such as direction, trajectory, orientation, and relative disp…

Vi-TacMan: Articulated Object Manipulation via Vision and Touch

2025-10-07 · Leiyao Cui, Zihang Zhao, Sirui Xie, Wenhuan Zhang 외 arxiv

Autonomous manipulation of articulated objects remains a fundamental challenge for robots in human environments. Vision-based methods can infer hidden kinematics but can yield imprecise estimates on unfamiliar objects. T…