paper-with-me

홈 › Papers

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

2026-06-07 · Yuan Zhang, Shiqi Zhang, Yedong Shen, Shuai Dong, Jiajun Deng, Xin Zhang, Yuxuan Gao, Jiajia Wu, Xin Nie, Zhiyuan Cheng, Jianmin Ji, Yanyong Zhang, Xingyi Zhang, Jia Pan arxiv

Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments. We argue that this stems from the lack of a unified geometry-aware manipulation representation, leaving existing VLAs vulnerable to low-level trajectory supervision, misaligned 3D features, and embodiment differences. To address this, we propose GEAR-VLA, a VLA framework for learning unified geometry-aware action representations for generalizable robotic manipulation. GEAR-VLA adopts coarse-to-fine action learning, where multi-source embodied pretraining equips the VLM with embodied reasoning and discrete action understanding before latent action tokens connect action semantics to a gradient-decoupled DiT continuous action expert. It further performs semantic-aligned 3D integration by aligning a trainable 3D spatial backbone with the VLA representation while freezing the original VLM-aligned visual pathway. To share this representation across robots, GEAR-VLA uses embodiment canonicalization, where embodiment-aware states and embodiment-invariant actions confine robot differences to the low-level interface. Extensive simulation and real-world experiments demonstrate strong generalization: GEAR-VLA achieves state-of-the-art performance on LIBERO, zero-shot LIBERO-Plus, and RoboTwin 2.0, reaches 85.9% success on AgileX and 81.0% on the pretraining-unseen LDT-01 embodiment, and obtains 90.1% success on a 6,360-trial universal grasping benchmark with 212 unseen objects. Code and models will be released at https://github.com/babynabeauty/GEAR-VLA.

📄 PDF Abstract BibTeX arXiv:2606.08530

Code (0)

등록된 구현이 없습니다.

Tasks

Action Understanding

Similar Papers 제목 키워드 기반

GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution

2023-07-17 · Yining Lu, Haoping Yu, Daniel Khashabi

Augmenting large language models (LLM) to use external tools enhances their performance across a variety of tasks. However, prior works over-rely on task-specific demonstration of tool use that limits their generalizabil…

Semantically Guided Representation Learning For Action Anticipation

2024-07-02 · Anxhelo Diko, Danilo Avola, Bardh Prenkaj, Federico Fontana 외

Action anticipation is the task of forecasting future activity from a partially observed sequence of events. However, this task is exposed to intrinsic future uncertainty and the difficulty of reasoning upon interconnect…

Action AnticipationRepresentation Learning

GEAR: GEometry-motion Alternating Refinement for Articulated Object Modeling with Gaussian Splatting

2026-04-09 · Jialin Li, Bin Fu, Ruiping Wang, Xilin Chen arxiv

High-fidelity interactive digital assets are essential for embodied intelligence and robotic interaction, yet articulated objects remain challenging to reconstruct due to their complex structures and coupled geometry-mot…

DYNAMO: Dependency-Aware Deep Learning Framework for Articulated Assembly Motion Prediction

2025-09-15 · Mayank Patel, Rahul Jain, Asim Unmesh, Karthik Ramani arxiv

Understanding the motion of articulated mechanical assemblies from static geometry remains a core challenge in 3D perception and design automation. Prior work on everyday articulated objects such as doors and laptops typ…

Point Clouds

Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models

2026-04-14 · Zijian Song, Qichang Li, Jiawei Zhou, Zhenlong Yuan 외 arxiv

At its core, robotic manipulation is a problem of vision-to-geometry mapping ($f(v) \rightarrow G$). Physical actions are fundamentally defined by geometric properties like 3D positions and spatial relationships. Consequ…

Zero-shot Generalization