paper-with-me

홈 › Papers

S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation

2026-06-26 · Zhipeng Xie, Zongyi Han, Xiangyi Wei, Shiliang Sun, Yang Li, Jing Zhao arxiv

Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, but their performance degrades significantly in long-horizon tasks due to cumulative error propagation. This limitation largely arises from static feature fusion mechanisms that rely on fixed weights to combine visual, language, and action representations, preventing the model from adapting to different phases of task execution. To address this limitation, we propose S$^2$-VLA, a framework that introduces a State-Space Guided Adaptive Attention (SSGAA) mechanism. SSGAA maintains a belief state that tracks task progression and generates dynamic gating weights to adaptively fuse information from three complementary sources visual features for spatial perception, task intents for high-level task planning, and temporal action sequences for execution consistency. This adaptive fusion allows the model to shift its focus throughout task execution, aligning with the evolving requirements of different task stages. Despite its compact 2B parameter size, S$^2$-VLA consistently outperforms larger 7B-scale models and achieves state-of-the-art performance on long-horizon manipulation benchmarks, including LIBERO and SimplerEnv. highlighting the importance of adaptive feature fusion for long-horizon robotic manipulation.

📄 PDF Abstract BibTeX arXiv:2606.27872

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hume: Introducing System-2 Thinking in Visual-Language-Action Model

2025-05-27 · Haoming Song, Delin Qu, Yuanqi Yao, Qizhi Chen 외

Humans practice slow thinking before performing actual actions when handling complex tasks in the physical world. This thinking paradigm, recently, has achieved remarkable advancement in boosting Large Language Models (L…

DenoisingVision-Language-Action

Learning with Language-Guided State Abstractions

2024-02-28 · Andi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers 외

We describe a framework for using natural language to design state abstractions for imitation learning. Generalizable policy learning in high-dimensional observation spaces is facilitated by well-designed state represent…

Imitation Learning

ReconVLA: An Uncertainty-Guided and Failure-Aware Vision-Language-Action Framework for Robotic Control

2026-04-17 · Lingling Chen, Zongyao Lyu, William J. Beksi arxiv

Vision-language-action (VLA) models have emerged as generalist robotic controllers capable of mapping visual observations and natural language instructions to continuous action sequences. However, VLAs provide no calibra…

Cross-Hand Latent Representation for Vision-Language-Action Models

2026-03-10 · Guangqi Jiang, Yutong Liang, Jianglong Ye, Jia-Yang Huang 외 arxiv

Dexterous manipulation is essential for real-world robot autonomy, mirroring the central role of human hand coordination in daily activity. Humans rely on rich multimodal perception--vision, sound, and language-guided in…

Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language Query

2019-10-01 · ICCV 2019 10 · Hao Wang, Cheng Deng, Junchi Yan, Dacheng Tao

Actor and action video segmentation from natural language query aims to selectively segment the actor and its action in a video based on an input textual description. Previous works mostly focus on learning simple correl…

Referring Expression SegmentationSegmentationVideo SegmentationVideo Semantic Segmentation