paper-with-me

홈 › Papers

Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models

2025-09-30 · Zhejia Cai, Yandan Yang, Xinyuan Chang, Shiyi Liang, Ronghan Chen, Feng Xiong, Mu Xu, Ruqi Huang arxiv

Latent Action Models (LAMs) enable Vision- Language-Action (VLA) systems to learn semantic action representations from large-scale unannotated data. Yet, we identify two bottlenecks of LAMs: 1) the commonly adopted end-to-end trained image encoder suffers from poor spatial understanding; 2) LAMs can be fragile when input frames are temporally distant, leading to limited temporal percep- tion. Such factors inevitably hinder stable and clear action modeling. To this end, we propose Farsighted-LAM, a latent action framework with geometry-aware spatial encoding and multi-scale temporal modeling, capturing structural priors and dynamic motion patterns from consecutive frames. We further propose SSM-VLA, an end-to-end VLA framework built upon Farsighted-LAM, which integrates structured perception with a visual Chain-of-Thought module to explicitly reason about environmental dynamics, enhancing decision consistency and interpretability. We validate SSM-VLA on multiple VLA tasks in both simulation and real-world settings, and achieve state-of- the-art performance. Our results demonstrate that our strategy of combining geometry-aware modeling, temporal coherence, and explicit reasoning is effective in enhancing the robustness and generalizability of embodied intelligence.

📄 PDF Abstract BibTeX arXiv:2509.26251

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Object Properties Inferring from and Transfer for Human Interaction Motions

2020-08-20 · Qian Zheng, Weikai Wu, Hanting Pan, Niloy Mitra 외

Humans regularly interact with their surrounding objects. Such interactions often result in strongly correlated motion between humans and the interacting objects. We thus ask: "Is it possible to infer object properties f…

Action RecognitionFine-grained Action RecognitionObject

Audio-driven Gesture Generation via Deviation Feature in the Latent Space

2025-03-27 · Jiahui Chen, Yang Huan, Runhua Shi, Chanfan Ding 외

Gestures are essential for enhancing co-speech communication, offering visual emphasis and complementing verbal interactions. While prior work has concentrated on point-level motion or fully supervised data-driven method…

Gesture GenerationVideo GenerationWeakly-supervised Learning

VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model

2026-02-10 · Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren 외 arxiv

Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevan…

Disentangled Neural Relational Inference for Interpretable Motion Prediction

2024-01-07 · Victoria M. Dax, Jiachen Li, Enna Sachdeva, Nakul Agarwal 외

Effective interaction modeling and behavior prediction of dynamic agents play a significant role in interactive motion planning for autonomous robots. Although existing methods have improved prediction accuracy, few rese…

Motion Planningmotion predictionPrediction

A structured latent space for human body motion generation

2021-06-07 · Mathieu Marsot, Stefanie Wuhrer, Jean-Sebastien Franco, Stephane Durocher

We propose a framework to learn a structured latent space to represent 4D human body motion, where each latent vector encodes a full motion of the whole 3D human shape. On one hand several data-driven skeletal animation …

3D geometryHuman motion predictionMotion Generationmotion prediction