paper-with-me

홈 › Papers

A Recipe for Generating 3D Worlds From a Single Image

2025-03-20 · Katja Schwarz, Denys Rozumnyi, Samuel Rota Bulò, Lorenzo Porzi, Peter Kontschieder

We introduce a recipe for generating immersive 3D worlds from a single image by framing the task as an in-context learning problem for 2D inpainting models. This approach requires minimal training and uses existing generative models. Our process involves two steps: generating coherent panoramas using a pre-trained diffusion model and lifting these into 3D with a metric depth estimator. We then fill unobserved regions by conditioning the inpainting model on rendered point clouds, requiring minimal fine-tuning. Tested on both synthetic and real images, our method produces high-quality 3D environments suitable for VR display. By explicitly modeling the 3D structure of the generated environment from the start, our approach consistently outperforms state-of-the-art, video synthesis-based methods along multiple quantitative image quality metrics. Project Page: https://katjaschwarz.github.io/worlds/

📄 PDF Abstract BibTeX arXiv:2503.16611

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

2024-08-20 · Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala 외

We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single tra…

Language ModelingLanguage Modelling

GILT: Generating Images from Long Text

2019-01-08 · Ori Bar El, Ori Licht, Netanel Yosephian

Creating an image reflecting the content of a long text is a complex process that requires a sense of creativity. For example, creating a book cover or a movie poster based on their summary or a food image based on its r…

Retrieval Augmented Recipe Generation

2024-11-13 · Guoshan Liu, Hailong Yin, Bin Zhu, Jingjing Chen 외

Given the potential applications of generating recipes from food images, this area has garnered significant attention from researchers in recent years. Existing works for recipe generation primarily utilize a two-stage t…

Recipe GenerationRetrieval

Structure-Aware Generation Network for Recipe Generation from Images

2020-09-02 · ECCV 2020 8 · Hao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan Miao

Sharing food has become very popular with the development of social media. For many real-world applications, people are keen to know the underlying recipes of a food item. In this paper, we are interested in automaticall…

Image CaptioningRecipe GenerationSentence

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

2025-07-29 · HunyuanWorld Team, Zhenwei Wang, Yuhao Liu, Junta Wu 외 arxiv

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods…