paper-with-me

Papers

VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation

2025-11-21 · Hanyu Zhou, Chuanhao Ma, Gim Hee Lee arxiv

Vision-language-action (VLA) models show potential for general robotic tasks, but remain challenging in spatiotemporally coherent manipulation, which requires fine-grained representations. Typically, existing methods embed 3D positions into visual representations to enhance the spatial precision of actions. However, these methods struggle to achieve temporally coherent control over action execution. In this work, we propose VLA-4D, a general VLA model with 4D awareness for spatiotemporally coherent robotic manipulation. Our model is guided by two key designs: 1) 4D-aware visual representation. We extract visual features, embed 1D time into 3D positions for 4D embeddings, and fuse them into a unified visual representation via a cross-attention mechanism. 2) Spatiotemporal action representation. We extend conventional spatial action representations with temporal information to enable the spatiotemporal planning, and align the multimodal representations into the LLM for spatiotemporal action prediction. Within this unified framework, the designed visual and action representations jointly make robotic manipulation spatially-smooth and temporally-coherent. In addition, we extend the VLA dataset with temporal action annotations for fine-tuning our model. Extensive experiments have been conducted to verify the superiority of our method across different tasks of robotic manipulation.

📄 PDF Abstract BibTeX arXiv:2511.17199

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

2025-08-12 · Lin Sun, Bin Xie, Yingfei Liu, Hao Shi 외 arxiv

Vision-Language-Action (VLA) models have emerged as a promising approach for enabling robots to follow language instructions and predict corresponding actions. However, current VLA models mainly rely on 2D visual inputs,…

Point Clouds

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models

2025-07-22 · Xiaoyan Wang, Zeju Li, Yifan Xu, Jiaxing Qi 외 arxiv

New era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on compressing holistic 3D scene informati…

Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models

2025-03-21 · Jianing Qi, Jiawei Liu, Hao Tang, Zhigang Zhu

Vision-Language Models (VLMs) excel at identifying and describing objects but struggle with spatial reasoning such as accurately understanding the relative positions of objects. Inspired by the dual-pathway (ventral-dors…

DiagnosticObject RecognitionSpatial Reasoning

LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

2024-09-26 · Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang 외

Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the developm…

3D Question Answering (3D-QA)PositionScene Understanding

Task-Aware Clustering for Prompting Vision-Language Models

2025-01-01 · CVPR 2025 1 · Fusheng Hao, Fengxiang He, Fuxiang Wu, Tichao Wang 외

Prompt learning has attracted widespread attention in adapting vision-language models to downstream tasks. Existing methods largely rely on optimization strategies to ensure the task-awareness of learnable prompts. D…

ClusteringDomain GeneralizationPrompt Learning