paper-with-me

홈 › Papers

Compositional Visual Generation and Inference with Energy Based Models

2020-04-13 · Yilun Du, Shuang Li, Igor Mordatch

A vital aspect of human intelligence is the ability to compose increasingly complex concepts out of simpler ideas, enabling both rapid learning and adaptation of knowledge. In this paper we show that energy-based models can exhibit this ability by directly combining probability distributions. Samples from the combined distribution correspond to compositions of concepts. For example, given a distribution for smiling faces, and another for male faces, we can combine them to generate smiling male faces. This allows us to generate natural images that simultaneously satisfy conjunctions, disjunctions, and negations of concepts. We evaluate compositional generation abilities of our model on the CelebA dataset of natural faces and synthetic 3D scene images. We also demonstrate other unique advantages of our model, such as the ability to continually learn and incorporate new concepts, or infer compositions of concept properties underlying an image.

📄 PDF Abstract BibTeX arXiv:2004.06030

Code (1)

yilundu/ebm_compositionality 공식 구현 tf

Similar Papers 제목 키워드 기반

Compositional Visual Generation with Composable Diffusion Models

2022-06-03 · Nan Liu, Shuang Li, Yilun Du, Antonio Torralba 외

Large text-guided diffusion models, such as DALLE-2, are able to generate stunning photorealistic images given natural language descriptions. While such models are highly flexible, they struggle to understand the composi…

Sentence

Robust and Controllable Object-Centric Learning through Energy-based Models

2022-10-11 · Ruixiang Zhang, Tong Che, Boris Ivanovic, Renhao Wang 외

Humans are remarkably good at understanding and reasoning about complex visual scenes. The capability to decompose low-level observations into discrete objects allows us to build a grounded abstract representation and id…

ObjectRepresentation LearningScene Generation

EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation

2026-05-12 · Sunung Mun, Sunghyun Cho, Jungseul Ok arxiv

Recent text-to-image (T2I) generators can synthesize realistic images, but still struggle with compositional prompts involving multiple objects, counts, attributes, and relations. We introduce EPIC (Efficient Predicate-G…

Text-to-Image Generation

Compositional Visual Generation with Energy Based Models

2020-12-01 · NeurIPS 2020 12 · Yilun Du, Shuang Li, Igor Mordatch

A vital aspect of human intelligence is the ability to compose increasingly complex concepts out of simpler ideas, enabling both rapid learning and adaptation of knowledge. In this paper we show that energy-based models …

Controllable and Compositional Generation with Latent-Space Energy-Based Models

2021-10-21 · NeurIPS 2021 12 · Weili Nie, Arash Vahdat, Anima Anandkumar

Controllable generation is one of the key requirements for successful adoption of deep generative models in real-world applications, but it still remains as a great challenge. In particular, the compositional ability to …

AttributeImage Generation