paper-with-me

홈 › Papers

Beyond World-Frame Action Heads: Motion-Centric Action Frames for Vision-Language-Action Models

2026-05-12 · Huoren Yang, Jianchao Zhao, Hu Yusong, Qiguan Ou, Yuyang Gao, Wei Ke, Yuhang He, SongLin Dong, Zhiheng Ma, Yihong Gong arxiv

Vision-Language-Action (VLA) models have advanced rapidly with stronger backbones, broader pre-training, and larger demonstration datasets, yet their action heads remain largely homogeneous: most directly predict action commands in a fixed world coordinate frame. We propose \textbf{MCF-Proto}, a lightweight action head that equips VLA policies with a Motion-Centric Action Frame (MCF) and a prototype-based action parameterization. At each step, the policy predicts a rotation $R_t \in SO(3)$, composes actions in the transformed local frame from a set of prototypes, and maps them back to the world frame for end-to-end training, using only standard demonstrations without auxiliary supervision. This simple design induces stable emergent structure. Without explicit directional labels, the learned local frames develop a stable geometric structure whose axes are strongly compatible with demonstrated end-effector motion. Meanwhile, actions in the learned representation become substantially more compact, with variation captured by fewer dominant directions and more regularly organized by shared prototypes. These structural properties translate into improved robustness, especially under geometric perturbations. Our results suggest that adding lightweight geometric and compositional structure to the action head can materially improve how VLA policies organize and generalize robotic manipulation behavior. An anonymized code repository is provided in the supplementary material.

📄 PDF Abstract BibTeX arXiv:2605.11809

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DirectSwap: Mask-Free Cross-Identity Training and Benchmarking for Expression-Consistent Video Head Swapping

2025-12-10 · Yanan Wang, Shengcai Liao, Panwen Hu, Xin Li 외 arxiv

Video head swapping aims to replace the entire head of a video subject, including facial identity, head shape, and hairstyle, with that of a reference image, while preserving the target body, background, and motion dynam…

Real-Time Simulated Avatar from Head-Mounted Sensors

2024-03-11 · CVPR 2024 1 · Zhengyi Luo, Jinkun Cao, Rawal Khirodkar, Alexander Winkler 외

We present SimXR, a method for controlling a simulated avatar from information (headset pose and cameras) obtained from AR / VR headsets. Due to the challenging viewpoint of head-mounted cameras, the human body is often …

Egocentric Pose EstimationHumanoid ControlPose Estimation

XR$^3$: An Extended Reality Platform for Social-Physical Human-Robot Interaction

2026-01-18 · Chao Wang, Anna Belardinelli, Michael Gienger arxiv

Social-physical human-robot interaction (spHRI) is difficult to study: building and programming robots that integrate multiple interaction modalities is costly and slow, while VR-based prototypes often lack physical cont…

Robotic VLA Benefits from Joint Learning with Motion Image Diffusion

2025-12-19 · Yu Fang, Kanchana Ranasinghe, Le Xue, Honglu Zhou 외 arxiv

Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they typically mimic expert trajectories wit…

OSM-Net: One-to-Many One-shot Talking Head Generation with Spontaneous Head Motions

2023-09-28 · Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai 외

One-shot talking head generation has no explicit head movement reference, thus it is difficult to generate talking heads with head motions. Some existing works only edit the mouth area and generate still talking heads, l…

Talking Head GenerationVideo Generation