paper-with-me

Papers

LucidDreaming: Controllable Object-Centric 3D Generation

2023-11-30 · Zhaoning Wang, Ming Li, Chen Chen

With the recent development of generative models, Text-to-3D generations have also seen significant growth, opening a door for creating video-game 3D assets from a more general public. Nonetheless, people without any professional 3D editing experience would find it hard to achieve precise control over the 3D generation, especially if there are multiple objects in the prompt, as using text to control often leads to missing objects and imprecise locations. In this paper, we present LucidDreaming as an effective pipeline capable of spatial and numerical control over 3D generation from only textual prompt commands or 3D bounding boxes. Specifically, our research demonstrates that Large Language Models (LLMs) possess 3D spatial awareness and can effectively translate textual 3D information into precise 3D bounding boxes. We leverage LLMs to get individual object information and their 3D bounding boxes as the initial step of our process. Then with the bounding boxes, We further propose clipped ray sampling and object-centric density blob bias to generate 3D objects aligning with the bounding boxes. We show that our method exhibits remarkable adaptability across a spectrum of mainstream Score Distillation Sampling-based 3D generation frameworks and our pipeline can even used to insert objects into an existing NeRF scene. Moreover, we also provide a dataset of prompts with 3D bounding boxes, benchmarking 3D spatial controllability. With extensive qualitative and quantitative experiments, we demonstrate that LucidDreaming achieves superior results in object placement precision and generation fidelity compared to current approaches, while maintaining flexibility and ease of use for non-expert users.

📄 PDF Abstract BibTeX arXiv:2312.00588

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationBenchmarkingLanguage ModellingLarge Language ModelNeRFObjectText to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

2026-08-24 · Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan 외 arxiv

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects i…

Novel View SynthesisSpatial Reasoning

CTRL-O: Language-Controllable Object-Centric Visual Representation Learning

2025-03-27 · CVPR 2025 1 · Aniket Didolkar, Andrii Zadaianchuk, Rabiul Awal, Maximilian Seitzer 외

Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current state-of-the-art object-centric models …

Image GenerationObjectObject DiscoveryQuestion Answering+4

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

2026-04-13 · Yuqian Yuan, Wenqiao Zhang, Juekai Lin, Yu Zhong 외 arxiv

Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasonin…

Scene UnderstandingSpatial Reasoning

EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses

2025-11-22 · Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy 외 arxiv

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controll…

Video GenerationVideo Prediction

Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation

2026-08-21 · Derui Li, Qian Qiao, Yuhao Sun, Wenhao Guo 외 arxiv

Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered…

Spatial ReasoningImage Generation