paper-with-me

홈 › Papers

RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation

2024-06-27 · Fanfan Liu, Feng Yan, Liming Zheng, Chengjian Feng, Yiyang Huang, Lin Ma

Utilizing Vision-Language Models (VLMs) for robotic manipulation represents a novel paradigm, aiming to enhance the model's ability to generalize to new objects and instructions. However, due to variations in camera specifications and mounting positions, existing methods exhibit significant performance disparities across different robotic platforms. To address this challenge, we propose RoboUniView in this paper, an innovative approach that decouples visual feature extraction from action learning. We first learn a unified view representation from multi-perspective views by pre-training on readily accessible data, and then derive actions from this unified view representation to control robotic manipulation. This unified view representation more accurately mirrors the physical world and is not constrained by the robotic platform's camera parameters. Thanks to this methodology, we achieve state-of-the-art performance on the demanding CALVIN benchmark, enhancing the success rate in the $D \to D$ setting from 93.0% to 96.2%, and in the $ABC \to D$ setting from 92.2% to 94.2%. Moreover, our model exhibits outstanding adaptability and flexibility: it maintains high performance under unseen camera parameters, can utilize multiple datasets with varying camera parameters, and is capable of joint cross-task learning across datasets. Code is provided for re-implementation. https://github.com/liufanfanlff/RoboUniview

📄 PDF Abstract BibTeX arXiv:2406.18977

Code (1)

liufanfanlff/robouniview 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingRobot ManipulationZero-shot Generalization

Similar Papers 제목 키워드 기반

RGB-Pointmap Pretraining for Unified 3D Scene Understanding

2026-04-02 · Ye Mao, Weixun Luo, Ranran Huang, Junpeng Jing 외 arxiv

Pretraining 3D encoders through alignment with Contrastive Language-Image Pre-training (CLIP) has emerged as a promising direction for learning generalizable representations for 3D scene understanding. In this paper, we …

Visual Question AnsweringRepresentation LearningScene ClassificationScene Understanding

VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking

2026-05-06 · Boyue Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu arxiv

UAV-ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two-stream methods suffer from isolated feature extraction and rely heavily on impli…

Visual Tracking

Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning

2026-02-24 · Haoyi Jiang, Liu Liu, Xinjie Wang, Yonghao He 외 arxiv

Vision-language models excel at 2D visual understanding but remain limited in 3D spatial reasoning. Existing approaches either depend on explicit 3D modalities, which limits scalability, or inject partial, view-condition…

Visual Reasoning

Unified Visual-Semantic Embeddings: Bridging Vision and Language With Structured Meaning Representations

2019-06-01 · CVPR 2019 6 · Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang 외

We propose the Unified Visual-Semantic Embeddings (Unified VSE) for learning a joint space of visual representation and textual semantics. The model unifies the embeddings of concepts at different levels: objects, attrib…

Contrastive LearningCross-Modal RetrievalRetrievalSentence

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation

2026-06-02 · Zeyuan Yang, Hao-Wei Chen, Xueyang Yu, Yuncong Yang 외 arxiv

Recent years have seen remarkable progress in unified vision-language models handling both multimodal understanding and generation within a single architecture. While autoregressive VLMs can reason across modalities, the…

multimodal generationImage GenerationText Generation