paper-with-me

홈 › Papers

Text2Scene: Generating Compositional Scenes from Textual Descriptions

2018-09-04 · CVPR 2019 6 · Fuwen Tan, Song Feng, Vicente Ordonez

In this paper, we propose Text2Scene, a model that generates various forms of compositional scene representations from natural language descriptions. Unlike recent works, our method does NOT use Generative Adversarial Networks (GANs). Text2Scene instead learns to sequentially generate objects and their attributes (location, size, appearance, etc) at every time step by attending to different parts of the input text and the current status of the generated scene. We show that under minor modifications, the proposed framework can handle the generation of different forms of scene representations, including cartoon-like scenes, object layouts corresponding to real images, and synthetic images. Our method is not only competitive when compared with state-of-the-art GAN-based methods using automatic metrics and superior based on human judgments but also has the advantage of producing interpretable results.

📄 PDF Abstract BibTeX arXiv:1809.01110

Code (3)

uvavision/Text2Scene 공식 구현 pytorch
123972/Deep-Learning-Final-Project
rafaelortegar/Deep-Learning-Final-Project

Similar Papers 제목 키워드 기반

ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment

2026-04-12 · Mingyu Dong, Chong Xia, Mingyuan Jia, Weichen Lyu 외 arxiv

Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstru…

3D Reconstruction

Frankenstein: Generating Semantic-Compositional 3D Scenes in One Tri-Plane

2024-03-24 · Han Yan, Yang Li, Zhennan Wu, Shenzhou Chen 외

We present Frankenstein, a diffusion-based framework that can generate semantic-compositional 3D scenes in a single pass. Unlike existing methods that output a single, unified 3D shape, Frankenstein simultaneously genera…

DenoisingObject Rearrangement

CG3D: Compositional Generation for Text-to-3D via Gaussian Splatting

2023-11-29 · Alexander Vilesov, Pradyumna Chari, Achuta Kadambi

With the onset of diffusion-based generative models and their ability to generate text-conditioned images, content generation has received a massive invigoration. Recently, these models have been shown to provide useful …

3D GenerationObjectText to 3D

MALeR: Improving Compositional Fidelity in Layout-Guided Generation

2025-11-08 · Shivank Saxena, Dhruv Srivastava, Makarand Tapaswi arxiv

Recent advances in text-to-image models have enabled a new era of creative and controllable image generation. However, generating compositional scenes with multiple subjects and attributes remains a significant challenge…

Image Generation

Semantic Score Distillation Sampling for Compositional Text-to-3D Generation

2024-10-11 · Ling Yang, Zixiang Zhang, Junlin Han, Bohan Zeng 외

Generating high-quality 3D assets from textual descriptions remains a pivotal challenge in computer graphics and vision research. Due to the scarcity of 3D data, state-of-the-art approaches utilize pre-trained 2D diffusi…

3D GenerationText to 3D