paper-with-me

홈 › Papers

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

2026-05-06 · Yihan Lin, Haoyang Li, Yang Li, Haitao Shen, Yihan Zhao, Chao Shao, Jing Zhang arxiv

Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and lack a systematic comparison. This work structures the study of latent action supervision from two perspectives: (i) regularizing the trajectory via image-based latent actions, and (ii) unifying the target space with action-based latent actions. Under a unified VLA baseline, we instantiate and compare four representative integration strategies. Our results reveal a formulation-task correspondence: image-based latent actions benefit long-horizon reasoning and scene-level generalization, whereas action-based latent actions excel at complex motor coordination. Furthermore, we find that directly supervising the VLM with discrete latent action tokens yields the most effective performance. Finally, our experiments offer initial insights into the benefits of latent action supervision in mixed-data, suggesting a promising direction for VLA training. Code is available at https://github.com/RUCKBReasoning/From_Pixels_to_Tokens.

📄 PDF Abstract BibTeX arXiv:2605.04678

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content

2026-07-17 · Ingo Ziegler, Martin Krebs, Desmond Elliott arxiv

Language models encode text as subword tokens, raw bytes, or rendered pixels, but these encodings are usually compared under modeling constraints that expose different amounts of linguistic content to models across diffe…

MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning

2025-11-27 · Kyeongha Rho, Hyeongkeun Lee, Jae Won Cho, Joon Son Chung arxiv

In this paper, we propose Mixture of Layer-Wise Tokens (MoLT), a parameter- and memory-efficient adaptation framework for audio-visual learning. The key idea of MoLT is to replace conventional, computationally heavy sequ…

audio-visual event localizationAudio-visual Question Answering

Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective

2025-11-20 · Jiahao Li, Yang Lu, Yachao Zhang, Yong Xie 외 arxiv

Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction cap…

Semantic Segmentation

FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

2025-06-04 · Xuanhua He, Quande Liu, Zixuan Ye, Weicai Ye 외

Fine-grained and efficient controllability on video diffusion transformers has raised increasing desires for the applicability. Recently, In-context Conditioning emerged as a powerful paradigm for unified conditional vid…

Video EditingVideo Generation

MaskBit: Embedding-free Image Generation via Bit Tokens

2024-09-24 · Mark Weber, Lijun Yu, Qihang Yu, Xueqing Deng 외

Masked transformer models for class-conditional image generation have become a compelling alternative to diffusion models. Typically comprising two stages - an initial VQGAN model for transitioning between latent space a…

Conditional Image GenerationImage GenerationImage Reconstruction