paper-with-me

홈 › Papers

BlobGAN: Spatially Disentangled Scene Representations

2022-05-05 · Dave Epstein, Taesung Park, Richard Zhang, Eli Shechtman, Alexei A. Efros

We propose an unsupervised, mid-level representation for a generative model of scenes. The representation is mid-level in that it is neither per-pixel nor per-image; rather, scenes are modeled as a collection of spatial, depth-ordered "blobs" of features. Blobs are differentiably placed onto a feature grid that is decoded into an image by a generative adversarial network. Due to the spatial uniformity of blobs and the locality inherent to convolution, our network learns to associate different blobs with different entities in a scene and to arrange these blobs to capture scene layout. We demonstrate this emergent behavior by showing that, despite training without any supervision, our method enables applications such as easy manipulation of objects within a scene (e.g., moving, removing, and restyling furniture), creation of feasible scenes given constraints (e.g., plausible rooms with drawers at a particular location), and parsing of real-world images into constituent parts. On a challenging multi-category dataset of indoor scenes, BlobGAN outperforms StyleGAN2 in image quality as measured by FID. See our project page for video results and interactive demo: https://www.dave.ml/blobgan

📄 PDF Abstract BibTeX arXiv:2205.02837

Code (0)

등록된 구현이 없습니다.

Tasks

Generative Adversarial Network

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Path Length Regularization 설명 없음
R1 Regularization R_INLINE_MATH_1 Regularization is a regularization technique and gradient penalty for training [generative adversarial…
HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
Weight Demodulation 설명 없음

Similar Papers 제목 키워드 기반

BlobGAN-3D: A Spatially-Disentangled 3D-Aware Generative Model for Indoor Scenes

2023-03-26 · Qian Wang, Yiqun Wang, Michael Birsak, Peter Wonka

3D-aware image synthesis has attracted increasing interest as it models the 3D nature of our real world. However, performing realistic object-level editing of the generated images in the multi-object scenario still remai…

3D-Aware Image SynthesisDisentanglementImage GenerationObject

Disentangled 3D Scene Generation with Layout Learning

2024-02-26 · Dave Epstein, Ben Poole, Ben Mildenhall, Alexei A. Efros 외

We introduce a method to generate 3D scenes that are disentangled into their component objects. This disentanglement is unsupervised, relying only on the knowledge of a large pretrained text-to-image model. Our key insig…

DisentanglementScene GenerationText to 3Dvalid

DEVIAS: Learning Disentangled Video Representations of Action and Scene

2023-11-30 · Kyungho Bae, Geo Ahn, Youngrae Kim, Jinwoo Choi

Video recognition models often learn scene-biased action representation due to the spurious correlation between actions and scenes in the training data. Such models show poor performance when the test data consists of vi…

Action RecognitionDecoderDisentanglementTemporal Action Localization+2

Physically Disentangled Representations

2022-04-11 · Tzofi Klinghoffer, Kushagra Tiwary, Arkadiusz Balata, Vivek Sharma 외

State-of-the-art methods in generative representation learning yield semantic disentanglement, but typically do not consider physical scene parameters, such as geometry, albedo, lighting, or camera. We posit that inverse…

AttributeClassificationDisentanglementEmotion Recognition+2

DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis

2022-12-22 · CVPR 2023 1 · Yinghao Xu, Menglei Chai, Zifan Shi, Sida Peng 외

Existing 3D-aware image synthesis approaches mainly focus on generating a single canonical object and show limited capacity in composing a complex scene containing a variety of objects. This work presents DisCoScene: a 3…

3D-Aware Image SynthesisImage GenerationObject