paper-with-me

Papers

Exploring Compositional Visual Generation with Latent Classifier Guidance

2023-04-25 · Changhao Shi, Haomiao Ni, Kai Li, Shaobo Han, Mingfu Liang, Martin Renqiang Min

Diffusion probabilistic models have achieved enormous success in the field of image generation and manipulation. In this paper, we explore a novel paradigm of using the diffusion model and classifier guidance in the latent semantic space for compositional visual tasks. Specifically, we train latent diffusion models and auxiliary latent classifiers to facilitate non-linear navigation of latent representation generation for any pre-trained generative model with a semantic latent space. We demonstrate that such conditional generation achieved by latent classifier guidance provably maximizes a lower bound of the conditional log probability during training. To maintain the original semantics during manipulation, we introduce a new guidance term, which we show is crucial for achieving compositionality. With additional assumptions, we show that the non-linear manipulation reduces to a simple latent arithmetic approach. We show that this paradigm based on latent classifier guidance is agnostic to pre-trained generative models, and present competitive results for both image generation and sequential manipulation of real and synthetic images. Our findings suggest that latent classifier guidance is a promising approach that merits further exploration, even in the presence of other strong competing methods.

📄 PDF Abstract BibTeX arXiv:2304.12536

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Compositional Video Generation via Inference-Time Guidance

2026-05-14 · Ariel Shaulov, Eitan Shaar, Amit Edenzon, Gal Chechik 외 arxiv

Text-to-video diffusion models generate realistic videos, but often fail on prompts requiring fine-grained compositional understanding, such as relations between entities, attributes, actions, and motion directions. We h…

Video Generation

Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models

2025-03-21 · CVPR 2025 1 · Davide Berasi, Matteo Farina, Massimiliano Mancini, Elisa Ricci 외

Vision-Language Models (VLMs) learn a shared feature space for text and images, enabling the comparison of inputs of different modalities. While prior works demonstrated that VLMs organize natural language representation…

Sky + Fire = Sunset. Exploring Parallels between Visually Grounded Metaphors and Image Classifiers

2020-07-01 · WS 2020 7 · Yuri Bizzoni, Simon Dobnik

This work explores the differences and similarities between neural image classifiers{'} mis-categorisations and visually grounded metaphors - that we could conceive as intentional mis-categorisations. We discuss the poss…

Controllable and Compositional Generation with Latent-Space Energy-Based Models

2021-10-21 · NeurIPS 2021 12 · Weili Nie, Arash Vahdat, Anima Anandkumar

Controllable generation is one of the key requirements for successful adoption of deep generative models in real-world applications, but it still remains as a great challenge. In particular, the compositional ability to …

AttributeImage Generation

Learning Graph Embeddings for Compositional Zero-shot Learning

2021-02-03 · CVPR 2021 1 · Muhammad Ferjad Naeem, Yongqin Xian, Federico Tombari, Zeynep Akata

In compositional zero-shot learning, the goal is to recognize unseen compositions (e.g. old dog) of observed visual primitives states (e.g. old, cute) and objects (e.g. car, dog) in the training set. This is challenging …

Compositional Zero-Shot LearningGraph EmbeddingTransfer LearningZero-Shot Learning