paper-with-me

홈 › Papers

Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models

2023-12-07 · Ivan Kapelyukh, Yifei Ren, Ignacio Alzugaray, Edward Johns

We introduce Dream2Real, a robotics framework which integrates vision-language models (VLMs) trained on 2D data into a 3D object rearrangement pipeline. This is achieved by the robot autonomously constructing a 3D representation of the scene, where objects can be rearranged virtually and an image of the resulting arrangement rendered. These renders are evaluated by a VLM, so that the arrangement which best satisfies the user instruction is selected and recreated in the real world with pick-and-place. This enables language-conditioned rearrangement to be performed zero-shot, without needing to collect a training dataset of example arrangements. Results on a series of real-world tasks show that this framework is robust to distractors, controllable by language, capable of understanding complex multi-object relations, and readily applicable to both tabletop and 6-DoF rearrangement tasks.

📄 PDF Abstract BibTeX arXiv:2312.04533

Code (0)

등록된 구현이 없습니다.

Tasks

Object Rearrangement

Similar Papers 제목 키워드 기반

PACA: Perspective-Aware Cross-Attention Representation for Zero-Shot Scene Rearrangement

2024-10-29 · Shutong Jin, Ruiyu Wang, Kuangyi Chen, Florian T. Pokorny

Scene rearrangement, like table tidying, is a challenging task in robotic manipulation due to the complexity of predicting diverse object arrangements. Web-scale trained generative models such as Stable Diffusion can aid…

Object

DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image

2025-03-13 · Qi Zhao, Zhan Ma, Pan Zhou

Recent developments in generative diffusion models have turned many dreams into realities. For video object insertion, existing methods typically require additional information, such as a reference video or a 3D asset of…

Object

Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow

2025-12-31 · Karthik Dharmarajan, Wenlong Huang, Jiajun Wu, Li Fei-Fei 외 arxiv

Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to translate such human-led motions into the l…

Reinforcement LearningVideo Generation

Adaptive Coordination in Social Embodied Rearrangement

2023-05-31 · Andrew Szot, Unnat Jain, Dhruv Batra, Zsolt Kira 외

We present the task of "Social Rearrangement", consisting of cooperative everyday tasks like setting up the dinner table, tidying a house or unpacking groceries in a simulated multi-agent environment. In Social Rearrange…

Diversity

D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement

2023-09-28 · YiXuan Wang, Mingtong Zhang, Zhuoran Li, Tarik Kelestemur 외

Scene representation is a crucial design choice in robotic manipulation systems. An ideal representation is expected to be 3D, dynamic, and semantic to meet the demands of diverse manipulation tasks. However, previous wo…