paper-with-me

홈 › Papers

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling

2026-06-19 · Tao Xu, Runhao Zhang, Zhijian Huang, Jiayi Guan, Jiaxin Wang, Yifan Ding, Yong-Lu Li, Long Chen, Guang Chen, Jinghui Lu arxiv

Occluded tasks remain a bottleneck in robot manipulation. Existing solutions either deploy additional physical cameras requiring training-inference camera parity, or rely on explicit 3D reconstruction with high computational cost. Moreover, both approaches rely on standard agent-view and wrist-view observations, while failing to capture occlusion information and future scene evolution. To this end, we propose UniviewVLA, a unified multiview Vision-Language-Action model with world modeling, which infers multiview scene evolution for action prediction from only standard two-camera observations. We demonstrate that by leveraging generated multiview future views from the world model, UniviewVLA reveals occluded cues and models future scene evolution, improving action prediction and removing the need for extra hardware or explicit reconstruction. Besides, to accelerate inference while preserving prediction accuracy, UniviewVLA develops Motion-Informative Token Compression, which compresses each generated view from 625 to 16 tokens and reduces per-view latency from 6-7s to 0.2-0.3s. UniviewVLA also proposes training-free Action-Entropy View Selection, which dynamically identifies the most action-informative view at different inference stages. Extensive experiments show that UniviewVLA achieves 95.8% on LIBERO and 4.60 on CALVIN ABCD to D, both standard occlusion-free benchmarks. On customized occlusion-focused tasks, it improves success rate from 40.0% to 73.3%, and average real-robot success rate by 33.4 points, demonstrating stronger occlusion-focused performance without sacrificing standard occlusion-free benchmarks.

📄 PDF Abstract BibTeX arXiv:2606.21501

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation3D Reconstruction

Results from the Paper

RankTaskDatasetModelMetrics
#1 Robot Manipulation CALVIN UniviewVLA avg. sequence length (D to D): 4.60

Similar Papers 제목 키워드 기반

Cross-Attentive Multiview Fusion of Vision-Language Embeddings

2026-04-14 · Tomas Berriel Martins, Martin R. Oswald, Javier Civera arxiv

Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, however, remains a challenging problem. Existing approaches typically…

2D Semantic Segmentation

Multiview Cauchy Estimator Feature Embedding for Depth and Inertial Sensor-Based Human Action Recognition

2016-08-07 · Yanan Guo, Lei LI, Weifeng Liu, Jun Cheng 외

The ever-growing popularity of Kinect and inertial sensors has prompted intensive research efforts on human action recognition. Since human actions can be characterized by multiple feature representations extracted from …

Action RecognitionTemporal Action Localization

Unsupervised Multiview Contrastive Language-Image Joint Learning with Pseudo-Labeled Prompts Via Vision-Language Model for 3D/4D Facial Expression Recognition

2025-05-14 · Muzammil Behzad

In this paper, we introduce MultiviewVLM, a vision-language model designed for unsupervised contrastive multiview representation learning of facial emotions from 3D/4D data. Our architecture integrates pseudo-labels deri…

Contrastive LearningFacial Expression RecognitionLanguage ModelingLanguage Modelling+1

Motubrain: An Advanced World Action Model for Robot Control

2026-04-30 · MotuBrain Team, Chendong Xiang, Fan Bao, Haitian Liu 외 arxiv

Vision-Language-Action (VLA) models generalize semantically well but often lack fine-grained modeling of world dynamics. We present Motubrain, a unified World Action Model that jointly models video and action under a Uni…

Video Generation

FUSER: Feed-Forward MUltiview 3D Registration Transformer and SE(3)$^N$ Diffusion Refinement

2025-12-10 · Haobo Jiang, Jin Xie, Jian Yang, Liang Yu 외 arxiv

Registration of multiview point clouds conventionally relies on extensive pairwise matching to build a pose graph for global synchronization, which is computationally expensive and inherently ill-posed without holistic g…

Computational EfficiencyPoint Clouds