paper-with-me

홈 › Papers

QuoVLA: Quotient Space for Vision-Language-Action Models

2026-05-24 · Xuan Wang, Yinan Wu, Haoran Duan, Jungong Han arxiv

Vision-Language-Action (VLA) models commonly adapt pretrained Vision-Language Models (VLMs) to robot control by mapping visual observations and language instructions to continuous actions. Existing approaches typically take an action-insufficiency view, assuming that pretrained VLM latents either lack directly usable action information or should be shielded from action-learning signals. Against this view, our \textit{Quotient Theory for VLA} shows that pretrained VLM latents are not action-insufficient but action-sufficient: they already contain the information needed for control, yet remain overcomplete by distinguishing prompt-level variations that induce the same optimal action behavior. To operationalize this theory, we propose QuoVLA, a quotient-space framework for VLA that compresses pretrained VLM latents into action-sufficient representations. Specifically, QuoVLA instantiates this principle with a quantization module and a dual-branch design with relative temporal-complexity regularization, preserving action-relevant information while removing prompt-level redundancy. Extensive experiments across multiple benchmarks demonstrate that QuoVLA achieves strong performance, with particularly notable improvements in generalization under visual, linguistic, and environmental distribution shifts. Our code will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2605.24890

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visual Language Hypothesis

2025-12-29 · Xiu Li arxiv

We study visual representation learning from a structural and topological perspective. We begin from a single hypothesis: that visual understanding presupposes a semantic language for vision, in which many perceptual obs…

Representation Learning

Manifold Learning in Quotient Spaces

2018-06-01 · CVPR 2018 6 · Éloi Mehr, André Lieutier, Fernando Sanchez Bermudez, Vincent Guitteny 외

When learning 3D shapes we are usually interested in their intrinsic geometry rather than in their orientation. To deal with the orientation variations the usual trick consists in augmenting the data to exhibit all possi…

Rapidly-Exploring Quotient-Space Trees: Motion Planning using Sequential Simplifications

2019-06-04 · Andreas Orthey, Marc Toussaint

Motion planning problems can be simplified by admissible projections of the configuration space to sequences of lower-dimensional quotient-spaces, called sequential simplifications. To exploit sequential simplifications,…

Motion Planning

The Urysohn Ladder: Recursive Metric Contraction for Scalable Continual Learning

2025-12-20 · Xin Li arxiv

Continual learning systems face a fundamental geometric obstacle: as experience accumulates on a fixed-capacity manifold, covering numbers grow linearly with time, eventually forcing representational overlap and catastro…

Continual Learning

A bundle framework for observer design on smooth manifolds with symmetry

2019-07-22 · Anant A. Joshi, D. H. S. Maithripala, Ravi N. Banavar

The article presents a bundle framework for nonlinear observer design on a manifold with a Lie group action. The group action on the manifold decomposes the manifold to a quotient structure and an orbit space, and the pr…