paper-with-me

홈 › Papers

ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models

2026-08-06 · Rishik Sathua, Haonan Chen, Katherine Driggs-Campbell arxiv

Large-scale visuomotor policies have demonstrated impressive performance across a wide range of robot manipulation tasks. However, despite this success, manipulation polices often entangle scene geometry with the corresponding viewpoint, learning where objects lie in an image rather than where it lies in the task space. This entanglement inherently limits the corresponding policy's ability to learn from viewpoint-diverse datasets (ex. DROID, BridgeV2) and generalize beyond the viewpoints captured in their training data. In this work, we present ARGUS, an observation pre-processing pipeline that uses large-scale 3D vision models to align image observations from arbitrary camera viewpoints into a canonical viewpoint before passing it to downstream visuomotor policies. Experiments across training datasets with varying levels of viewpoint diversity, from fixed multi-view camera configurations to highly varied camera placements, show that our method consistently outperforms prior approaches across both limited-view and view-diverse training regimes. In efficiency comparisons, ARGUS demonstrates an ability to learn from view-diverse data, converging to high success rates 4-6x faster than previous methods by leveraging a simplified observation space. Overall, our findings show that leveraging large-scale 3D vision models reduces the learning burden on visuomotor policies, enabling more efficient learning from large-scale, viewpoint-diverse robot datasets.

📄 PDF Abstract BibTeX arXiv:2608.05579

Code (2)

BaiShuanghao/my_arXiv_daily ★ 208
arxivsub/arXivSub_daily_arxiv ★ 4

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

Argus: Smartphone-enabled Human Cooperation via Multi-Agent Reinforcement Learning for Disaster Situational Awareness

2019-04-29 · Vidyasagar Sadhu, Gabriel Salles-Loustau, Dario Pompili, Saman Zonouz 외

Argus exploits a Multi-Agent Reinforcement Learning (MARL) framework to create a 3D mapping of the disaster scene using agents present around the incident zone to facilitate the rescue operations. The agents can be both …

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Extreme dynamic symmetry enables omnidirectional and multifunctional robots

2026-05-28 · Jiaxun Liu, Boxi Xia, Boyuan Chen arxiv

Symmetry is a central organizing principle in natural systems, yet its use as a unifying design strategy in robotics has largely remained limited to geometric form. We show that symmetry can instead be leveraged at the l…

Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes

2026-06-29 · Xi Li, Linyuan Li, Yan Wu, Tong Rao 외 arxiv

Metric feed-forward 3D reconstruction for panoramic data remains under-explored due to the lack of large-scale panoramic RGB-D training data. We present Realsee3D, a hybrid dataset of 10K indoor scenes (1K real, 9K synth…

Camera Pose EstimationMulti-Task Learning3D ReconstructionDepth Estimation

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models

2025-07-17 · Yifan Xu, Chao Zhang, Hanqi Jiang, Xiaoyan Wang 외

Advancements in foundation models have made it possible to conduct applications in various downstream tasks. Especially, the new era has witnessed a remarkable capability to extend Large Language Models (LLMs) for tackli…

3D Point Cloud ReconstructionPoint cloud reconstructionScene Understanding

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

2026-07-28 · Weiming Zhuang, Jiabo Huang, Jingtao Li, Zhizhong Li 외 arxiv

Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two …