paper-with-me

홈 › Papers

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

2026-02-23 · Manish Kumar Govind, Dominick Reilly, Pu Wang, Srijan Das arxiv

Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-action (VLA) models without explicit robot action supervision. However, latent actions derived solely from RGB observations primarily encode appearance-driven dynamics and lack explicit 3D geometric structure, which is essential for precise and contact-rich manipulation. To address this limitation, we introduce UniLACT, a transformer-based VLA model that incorporates geometric structure through depth-aware latent pretraining, enabling downstream policies to inherit stronger spatial priors. To facilitate this process, we propose UniLARN, a unified latent action learning framework based on inverse and forward dynamics objectives that learns a shared embedding space for RGB and depth while explicitly modeling their cross-modal interactions. This formulation produces modality-specific and unified latent action representations that serve as pseudo-labels for the depth-aware pretraining of UniLACT. Extensive experiments in both simulation and real-world settings demonstrate the effectiveness of depth-aware unified latent action representations. UniLACT consistently outperforms RGB-based latent action baselines under in-domain and out-of-domain pretraining regimes, as well as on both seen and unseen manipulation tasks.The project page is at https://manishgovind.github.io/unilact-vla/

📄 PDF Abstract BibTeX arXiv:2602.20231

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

2025-10-16 · Yixuan Li, Yuhui Chen, Mingcai Zhou, Haoran Li 외 arxiv

Spatial perception and reasoning are crucial for Vision-Language-Action (VLA) models to accomplish fine-grained manipulation tasks. However, existing approaches often lack the ability to understand and reason over the es…

Spatial Reasoning

GLaD: Geometric Latent Distillation for Vision-Language-Action Models

2025-12-10 · Minghao Guo, Meng Cao, Jiachen Tao, Rongtao Xu 외 arxiv

Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we introduce GLaD, a geometry-aware VLA fra…

Knowledge DistillationSpatial Reasoning

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action

2026-06-04 · Boyang Zhang, Lianlei Shan arxiv

Vision-Language-Action (VLA) policies remain brittle in long-horizon control, where one-pass action decoding offers limited inference-time deliberation. Explicit chain-of-thought adds reasoning depth but incurs token-gen…

DALE-CT: Depth-Aware Foundation Models for Computed Tomography

2026-06-05 · Evan W. Damron, Mahmut S. Gokmen, Mitchell A. Klusty, Caroline N. Leach 외 arxiv

Recent breakthroughs in self-supervised learning (SSL), such as the Latent-Euclidean Joint-Embedding Predictive Architecture (LeJEPA), alongside successes in integrating visual encoders with language models, have driven …

Multiple Instance LearningSelf-Supervised Learning

Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning

2023-07-27 · ICCV 2023 1 · Junwen He, Yifan Wang, Lijun Wang, Huchuan Lu 외

Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this ch…

Depth EstimationPanoptic SegmentationSegmentation