paper-with-me

홈 › Papers

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models

2025-07-31 · Xiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang, Kaixin Wang, Yanjiang Guo, Rushuai Yang, Yucen Wang, Xinquan Xiao, Li Zhao, Jianyu Chen, Jiang Bian arxiv

Vision-Language-Action (VLA) models have emerged as a popular paradigm for learning robot manipulation policies that can follow language instructions and generalize to novel scenarios. Recent works have begun to explore the incorporation of latent actions, abstract representations of motion between two frames, into VLA pre-training. In this paper, we introduce villa-X, a novel Vision-Language-Latent-Action (ViLLA) framework that advances latent action modeling for learning generalizable robot manipulation policies. Our approach improves both how latent actions are learned and how they are incorporated into VLA pre-training. We demonstrate that villa-X can generate latent action plans in a zero-shot fashion, even for unseen embodiments and open-vocabulary symbolic understanding. This capability enables villa-X to achieve superior performance across diverse simulation tasks in SIMPLER and on two real-world robotic setups involving both gripper and dexterous hand manipulation. These results establish villa-X as a principled and scalable paradigm for learning generalizable robot manipulation policies. We believe it provides a strong foundation for future research.

📄 PDF Abstract BibTeX arXiv:2507.23682

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

Latent Action Pretraining Through World Modeling

2025-09-22 · Bahey Tharwat, Yara Nasser, Ali Abouzeid, Ian Reid arxiv

Vision-Language-Action (VLA) models have gained popularity for learning robotic manipulation tasks that follow language instructions. State-of-the-art VLAs, such as OpenVLA and $π_{0}$, were trained on large-scale, manua…

Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models

2025-09-30 · Zhejia Cai, Yandan Yang, Xinyuan Chang, Shiyi Liang 외 arxiv

Latent Action Models (LAMs) enable Vision- Language-Action (VLA) systems to learn semantic action representations from large-scale unannotated data. Yet, we identify two bottlenecks of LAMs: 1) the commonly adopted end-t…

UV-SAM: Adapting Segment Anything Model for Urban Village Identification

2024-01-16 · Xin Zhang, Yu Liu, Yuming Lin, Qingmin Liao 외

Urban villages, defined as informal residential areas in or around urban centers, are characterized by inadequate infrastructures and poor living conditions, closely related to the Sustainable Development Goals (SDGs) on…

image-classificationImage ClassificationSemantic Segmentation

Village-Net Clustering: A Rapid approach to Non-linear Unsupervised Clustering of High-Dimensional Data

2025-01-16 · Aditya Ballal, Esha Datta, Gregory A. DePaul, Erik Carlsson 외

Clustering large high-dimensional datasets with diverse variable is essential for extracting high-level latent information from these datasets. Here, we developed an unsupervised clustering algorithm, we call "Village-Ne…

BenchmarkingClusteringCommunity Detection

VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration

2026-02-04 · Jaeyoon Jung, Yejun Yoon, Kunwoo Park arxiv

This paper describes VILLAIN, a multimodal fact-checking system that verifies image-text claims through prompt-based multi-agent collaboration. For the AVerImaTeC shared task, VILLAIN employs vision-language model agents…