paper-with-me

홈 › Papers

Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation

2023-07-27 · William Shen, Ge Yang, Alan Yu, Jansen Wong, Leslie Pack Kaelbling, Phillip Isola

Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often lacking in 2D image features. This work bridges this 2D-to-3D gap for robotic manipulation by leveraging distilled feature fields to combine accurate 3D geometry with rich semantics from 2D foundation models. We present a few-shot learning method for 6-DOF grasping and placing that harnesses these strong spatial and semantic priors to achieve in-the-wild generalization to unseen objects. Using features distilled from a vision-language model, CLIP, we present a way to designate novel objects for manipulation via free-text natural language, and demonstrate its ability to generalize to unseen expressions and novel categories of objects.

📄 PDF Abstract BibTeX arXiv:2308.07931

Code (1)

f3rm/f3rm 공식 구현 jax

Tasks

3D geometryFew-Shot LearningLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

LensDFF: Language-enhanced Sparse Feature Distillation for Efficient Few-Shot Dexterous Manipulation

2025-03-05 · Qian Feng, David S. Martinez Lema, Jianxiang Feng, Zhaopeng Chen 외

Learning dexterous manipulation from few-shot demonstrations is a significant yet challenging problem for advanced, human-like robotic systems. Dense distilled feature fields have addressed this challenge by distilling r…

NeRFNeural Rendering

Vector Quantized Feature Fields for Fast 3D Semantic Lifting

2025-03-09 · George Tang, Aditya Agarwal, Weiqiao Han, Trevor Darrell 외

We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding multiscale pixel-aligned feature maps, which…

Embodied Question AnsweringQuestion Answering

Neural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping

2024-10-30 · Qianxu Wang, Congyue Deng, Tyler Ga Wei Lum, Yuanpei Chen 외

One-shot transfer of dexterous grasps to novel scenes with object and context variations has been a challenging problem. While distilled feature fields from large vision models have enabled semantic correspondences acros…

Decoder

LERF: Language Embedded Radiance Fields

2023-03-16 · ICCV 2023 1 · Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa 외

Humans describe the physical world using natural language to refer to specific 3D locations based on a vast range of properties: visual appearance, semantics, abstract associations, or actionable affordances. In this wor…

NeRF

Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields

2025-10-03 · Zhiting Mei, Ola Shorinwa, Anirudha Majumdar arxiv

Semantic distillation in radiance fields has spurred significant advances in open-vocabulary robot policies, e.g., in manipulation and navigation, founded on pretrained semantics from large vision models. While prior wor…

Object LocalizationPose Estimation