paper-with-me

홈 › Papers

Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D Inpainting

2024-11-14 · Yian Wang, Xiaowen Qiu, Jiageng Liu, Zhehuan Chen, Jiting Cai, YuFei Wang, Tsun-Hsuan Wang, Zhou Xian, Chuang Gan

Creating large-scale interactive 3D environments is essential for the development of Robotics and Embodied AI research. Current methods, including manual design, procedural generation, diffusion-based scene generation, and large language model (LLM) guided scene design, are hindered by limitations such as excessive human effort, reliance on predefined rules or training datasets, and limited 3D spatial reasoning ability. Since pre-trained 2D image generative models better capture scene and object configuration than LLMs, we address these challenges by introducing Architect, a generative framework that creates complex and realistic 3D embodied environments leveraging diffusion-based 2D image inpainting. In detail, we utilize foundation visual perception models to obtain each generated object from the image and leverage pre-trained depth estimation models to lift the generated 2D image to 3D space. Our pipeline is further extended to a hierarchical and iterative inpainting process to continuously generate placement of large furniture and small objects to enrich the scene. This iterative structure brings the flexibility for our method to generate or refine scenes from various starting points, such as text, floor plans, or pre-arranged environments.

📄 PDF Abstract BibTeX arXiv:2411.09823

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationImage InpaintingImage to 3DLanguage ModelingLanguage ModellingLarge Language ModelScene GenerationSpatial Reasoning

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

Demonstrating ViviDoc: Generating Interactive Documents through Human-Agent Collaboration

2026-03-02 · Yinghao Tang, Yupeng Xie, Yingchaojie Feng, Tingfeng Lan 외 arxiv

Interactive articles help readers engage with complex ideas through exploration, yet creating them remains costly, requiring both domain expertise and web development skills. Recent LLM-based agents can automate content …

VividDream: Generating 3D Scene with Ambient Dynamics

2024-05-30 · Yao-Chih Lee, Yi-Ting Chen, Andrew Wang, Ting-Hsuan Liao 외

We introduce VividDream, a method for generating explorable 4D scenes with ambient dynamics from a single input image or text prompt. VividDream first expands an input image into a static 3D point cloud through iterative…

ViviDoc: Generating Interactive Documents through Human-Agent Collaboration

2026-03-30 · Yinghao Tang, Yupeng Xie, Yingchaojie Feng, Tingfeng Lan 외 arxiv

Interactive documents help readers engage with complex ideas through dynamic visualization, interactive animations, and exploratory interfaces. However, creating such documents remains costly, as it requires both domain …

VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis

2026-02-01 · Chengyuan Ma, Jiawei Jin, Ruijie Xiong, Chunxiang Jin 외 arxiv

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the r…

Speech Synthesis

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

2024-11-22 · Jiahao Hu, Tianxiong Zhong, Xuebo Wang, Boyuan Jiang 외

Diffusion-based image editing models have made remarkable progress in recent years. However, achieving high-quality video editing remains a significant challenge. One major hurdle is the absence of open-source, large-sca…

Video Editing