paper-with-me

Papers

3D Neural Scene Representations for Visuomotor Control

2021-07-08 · Yunzhu Li, Shuang Li, Vincent Sitzmann, Pulkit Agrawal, Antonio Torralba

Humans have a strong intuitive understanding of the 3D environment around us. The mental model of the physics in our brain applies to objects of different materials and enables us to perform a wide range of manipulation tasks that are far beyond the reach of current robots. In this work, we desire to learn models for dynamic 3D scenes purely from 2D visual observations. Our model combines Neural Radiance Fields (NeRF) and time contrastive learning with an autoencoding framework, which learns viewpoint-invariant 3D-aware scene representations. We show that a dynamics model, constructed over the learned representation space, enables visuomotor control for challenging manipulation tasks involving both rigid bodies and fluids, where the target is specified in a viewpoint different from what the robot operates on. When coupled with an auto-decoding framework, it can even support goal specification from camera viewpoints that are outside the training distribution. We further demonstrate the richness of the learned 3D dynamics model by performing future prediction and novel view synthesis. Finally, we provide detailed ablation studies regarding different system designs and qualitative analysis of the learned representations.

📄 PDF Abstract BibTeX arXiv:2107.04004

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningFuture predictionNeRFNovel View Synthesis

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
Robinhood Customer Care Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Unified Visuomotor Targets: Supervising VLAs Beyond Physical Actions

2026-08-04 · Zhenyang Feng, Unnat Jain arxiv

VLA models are trained to predict robot actions from visual and language observations. This is a natural choice, but it creates a mismatch: VLMs encode rich, high-level representations of scenes and goals, while robot ac…

SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Planning and Control

2017-10-02 · Arunkumar Byravan, Felix Leeb, Franziska Meier, Dieter Fox

In this work, we present an approach to deep visuomotor control using structured deep dynamics models. Our deep dynamics model, a variant of SE3-Nets, learns a low-dimensional pose embedding for visuomotor control via an…

Decoder

Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues

2025-11-13 · Nikolaos Tsagkas, Andreas Sochopoulos, Duolikun Danier, Sethu Vijayakumar 외 arxiv

The adoption of pre-trained visual representations (PVRs), leveraging features from large-scale vision models, has become a popular paradigm for training visuomotor policies. However, these powerful representations can e…

Universal Planning Networks: Learning Generalizable Representations for Visuomotor Control

2018-07-01 · ICML 2018 7 · Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine 외

A key challenge in complex visuomotor control is learning abstract representations that are effective for specifying goals, planning, and generalization. To this end, we introduce universal planning networks (UPN). …

Imitation LearningReinforcement Learning

Generalization Through Hand-Eye Coordination: An Action Space for Learning Spatially-Invariant Visuomotor Control

2021-02-28 · Chen Wang, Rui Wang, Ajay Mandlekar, Li Fei-Fei 외

Imitation Learning (IL) is an effective framework to learn visuomotor skills from offline demonstration data. However, IL methods often fail to generalize to new scene configurations not covered by training data. On the …

Imitation LearningZero-shot Generalization