paper-with-me

홈 › Papers

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

2025-08-28 · Wei Li, Renshan Zhang, Rui Shao, Jie He, Liqiang Nie arxiv

Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high computational overhead that limits scalability and deployment.We propose CogVLA, a Cognition-Aligned Vision-Language-Action framework that leverages instruction-driven routing and sparsification to improve both efficiency and performance. CogVLA draws inspiration from human multimodal coordination and introduces a 3-stage progressive architecture. 1) Encoder-FiLM based Aggregation Routing (EFA-Routing) injects instruction information into the vision encoder to selectively aggregate and compress dual-stream visual tokens, forming a instruction-aware latent representation. 2) Building upon this compact visual encoding, LLM-FiLM based Pruning Routing (LFP-Routing) introduces action intent into the language model by pruning instruction-irrelevant visually grounded tokens, thereby achieving token-level sparsity. 3) To ensure that compressed perception inputs can still support accurate and coherent action generation, we introduce V-L-A Coupled Attention (CAtten), which combines causal vision-language attention with bidirectional action parallel decoding. Extensive experiments on the LIBERO benchmark and real-world robotic tasks demonstrate that CogVLA achieves state-of-the-art performance with success rates of 97.4% and 70.0%, respectively, while reducing training costs by 2.5-fold and decreasing inference latency by 2.8-fold compared to OpenVLA. CogVLA is open-sourced and publicly available at https://github.com/JiuTian-VL/CogVLA.

📄 PDF Abstract BibTeX arXiv:2508.21046

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Aligned Image-Word Representations Improve Inductive Transfer Across Vision-Language Tasks

2017-04-02 · ICCV 2017 10 · Tanmay Gupta, Kevin Shih, Saurabh Singh, Derek Hoiem

An important goal of computer vision is to build systems that learn visual representations over time that can be applied to many tasks. In this paper, we investigate a vision-language embedding as a core representation a…

Multi-Task LearningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment

2025-06-28 · Amir Aghdam, Vincent Tao Hu

We address the task of zero-shot fine-grained video classification, where no video examples or temporal annotations are available for unseen action classes. While contrastive vision-language models such as SigLIP demonst…

Dynamic Time WarpingLarge Language ModelOpen Set Learningtext similarity+2

Transformers in Action Recognition: A Review on Temporal Modeling

2022-12-29 · Elham Shabaninia, Hossein Nezamabadi-pour, Fatemeh Shafizadegan

In vision-based action recognition, spatio-temporal features from different modalities are used for recognizing activities. Temporal modeling is a long challenge of action recognition. However, there are limited methods …

Action Recognition

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model

2026-05-29 · Xiang Zhu, Puzhen Yuan, Yichen Liu, Jianyu Chen arxiv

Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions. While latent…

Representation Learning

Rescaling Egocentric Vision

2020-06-23 · Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari 외

This paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS. The effort culminates in EPIC-KITCHENS-100, a collection of 100 hours, 20M frames, 90K actions in 700 variable-length …

Action AnticipationAction DetectionAction RecognitionCross-Modal Retrieval+3