paper-with-me

홈 › Papers

Towards Controllable Image Generation through Representation-Conditioned Diffusion Models

2026-05-26 · Nithesh Chandher Karthikeyan, Jonas Unger, Gabriel Eilertsen arxiv

Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text prompts or semantic maps, which require extensively annotated datasets. In this preliminary work, we explore diffusion models conditioned on representations from a pre-trained self-supervised model. The self-conditioning mechanism not only improves the quality of unconditional image generation, but also provides a representation space that can be used to control the generation. We explore this conditioning space by identifying directions of variations, and demonstrate promising properties in terms of smoothness and disentanglement.

📄 PDF Abstract BibTeX arXiv:2605.27343

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models

2025-12-03 · Shaoheng Fang, Chaohui Yu, Fan Wang, Qixing Huang arxiv

We introduce MVRoom, a controllable novel view synthesis (NVS) pipeline for 3D indoor scenes that uses multi-view diffusion conditioned on a coarse 3D layout. MVRoom employs a two-stage design in which the 3D layout is u…

Novel View SynthesisScene Generation

Multiverse: Language-Conditioned Multi-Game Level Blending via Shared Representation

2026-03-25 · In-Chang Baek, Jiyun Jung, Geum-Hwan Hwang, Sung-Hyun Kim 외 arxiv

Text-to-level generation aims to translate natural language descriptions into structured game levels, enabling intuitive control over procedural content generation. While prior text-to-level generators are typically limi…

From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation

2024-04-23 · Zehuan Huang, Hongxing Fan, Lipeng Wang, Lu Sheng

Recent advancements in controllable human image generation have led to zero-shot generation using structural signals (e.g., pose, depth) or facial appearance. Yet, generating human images conditioned on multiple parts of…

Image Generation

GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation

2026-08-17 · Tianchen Deng, Xuefeng Chen, Shuang Wu, Qu Chen 외 arxiv

Driving World Models (DWMs) have recently advanced rapidly with generative models, yet most existing methods mainly focus on conditional scene generation and lack explicit 3D scene understanding, language-grounded reason…

Scene UnderstandingScene GenerationVisual Grounding

Chord-Conditioned Melody Harmonization with Controllable Harmonicity

2022-02-17 · Shangda Wu, Xiaobing Li, Maosong Sun

Melody harmonization has long been closely associated with chorales composed by Johann Sebastian Bach. Previous works rarely emphasised chorale generation conditioned on chord progressions, and there has been a lack of f…