paper-with-me

홈 › Papers

One-Shot Neural Fields for 3D Object Understanding

2022-10-21 · Valts Blukis, Taeyeop Lee, Jonathan Tremblay, Bowen Wen, In So Kweon, Kuk-Jin Yoon, Dieter Fox, Stan Birchfield

We present a unified and compact scene representation for robotics, where each object in the scene is depicted by a latent code capturing geometry and appearance. This representation can be decoded for various tasks such as novel view rendering, 3D reconstruction (e.g. recovering depth, point clouds, or voxel maps), collision checking, and stable grasp prediction. We build our representation from a single RGB input image at test time by leveraging recent advances in Neural Radiance Fields (NeRF) that learn category-level priors on large multiview datasets, then fine-tune on novel objects from one or few views. We expand the NeRF model for additional grasp outputs and explore ways to leverage this representation for robotics. At test-time, we build the representation from a single RGB input image observing the scene from only one viewpoint. We find that the recovered representation allows rendering from novel views, including of occluded object parts, and also for predicting successful stable grasps. Grasp poses can be directly decoded from our latent representation with an implicit grasp decoder. We experimented in both simulation and real world and demonstrated the capability for robust robotic grasping using such compact representation. Website: https://nerfgrasp.github.io

📄 PDF Abstract BibTeX arXiv:2210.12126

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionDecoderNeRFObjectPose PredictionRobotic Grasping

Similar Papers 제목 키워드 기반

OpenObj: Open-Vocabulary Object-Level Neural Radiance Fields with Fine-Grained Understanding

2024-06-12 · Yinan Deng, Jiahui Wang, Jingyu Zhao, Jianyu Dou 외

In recent years, there has been a surge of interest in open-vocabulary 3D scene reconstruction facilitated by visual language models (VLMs), which showcase remarkable capabilities in open-set retrieval. However, existing…

3D Scene ReconstructionNeRFObjectRetrieval+2

Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation

2023-07-27 · William Shen, Ge Yang, Alan Yu, Jansen Wong 외

Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often …

3D geometryFew-Shot LearningLanguage ModelingLanguage Modelling

OpenOcc: Open Vocabulary 3D Scene Reconstruction via Occupancy Representation

2024-03-18 · Haochen Jiang, Yueming Xu, Yihan Zeng, Hang Xu 외

3D reconstruction has been widely used in autonomous navigation fields of mobile robotics. However, the former research can only provide the basic geometry structure without the capability of open-world scene understandi…

3D Reconstruction3D Scene ReconstructionAutonomous NavigationScene Understanding+1

ProfileFoundry: A Synthetic Person-Object Substrate for Privacy, Memory, and Tool-Use Evaluation in LLM Agent

2026-06-24 · Sriram Selvam, Anneswa Ghosh arxiv

Foundation-model research increasingly needs data about people: user state, personal histories, relationships, contact-like fields, documents, and longitudinal updates. Real user data is difficult to share, perturb, audi…

Zero-shot counting with a dual-stream neural network model

2024-05-16 · Jessica A. F. Thompson, Hannah Sheahan, Tsvetomira Dumbalska, Julian Sandbrink 외

Deep neural networks have provided a computational framework for understanding object recognition, grounded in the neurophysiology of the primate ventral stream, but fail to account for how we process relational aspects …

Object RecognitionZero-Shot Counting