paper-with-me

Papers

Collaborative Diffusion for Multi-Modal Face Generation and Editing

2023-04-20 · CVPR 2023 1 · Ziqi Huang, Kelvin C. K. Chan, Yuming Jiang, Ziwei Liu

Diffusion models arise as a powerful generative tool recently. Despite the great progress, existing diffusion models mainly focus on uni-modal control, i.e., the diffusion process is driven by only one modality of condition. To further unleash the users' creativity, it is desirable for the model to be controllable by multiple modalities simultaneously, e.g., generating and editing faces by describing the age (text-driven) while drawing the face shape (mask-driven). In this work, we present Collaborative Diffusion, where pre-trained uni-modal diffusion models collaborate to achieve multi-modal face generation and editing without re-training. Our key insight is that diffusion models driven by different modalities are inherently complementary regarding the latent denoising steps, where bilateral connections can be established upon. Specifically, we propose dynamic diffuser, a meta-network that adaptively hallucinates multi-modal denoising steps by predicting the spatial-temporal influence functions for each pre-trained uni-modal model. Collaborative Diffusion not only collaborates generation capabilities from uni-modal diffusion models, but also integrates multiple uni-modal manipulations to perform multi-modal editing. Extensive qualitative and quantitative experiments demonstrate the superiority of our framework in both image quality and condition consistency.

📄 PDF Abstract BibTeX arXiv:2304.10530

Code (1)

ziqihuangg/collaborative-diffusion 공식 구현 pytorch

Tasks

DenoisingFace Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

2025-11-16 · Yushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou 외 arxiv

While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, …

Reasoning with Autoregressive-Diffusion Collaborative Thoughts

2026-02-02 · Mu Yuan, Liekang Zeng, Guoliang Xing, Lan Zhang 외 arxiv

Autoregressive and diffusion models represent two complementary generative paradigms. Autoregressive models excel at sequential planning and constraint composition, yet struggle with tasks that require explicit spatial o…

Question AnsweringSpatial Reasoning

Variational Test-time Optimization for Diffusion Synchronization

2026-06-14 · Hyunsoo Lee, Farrin Marouf Sofian, Kushagra Pandey, Stephan Mandt arxiv

Collaborative generation, which coordinates multiple diffusion trajectories to extend the capabilities of pretrained priors, has emerged as a powerful paradigm for extending the applicability of diffusion models. Among e…

Multi-Agent Amodal Completion: Direct Synthesis with Fine-Grained Semantic Guidance

2025-09-22 · Hongxing Fan, Lipeng Wang, Haohua Chen, Zehuan Huang 외 arxiv

Amodal completion, generating invisible parts of occluded objects, is vital for applications like image editing and AR. Prior methods face challenges with data needs, generalization, or error accumulation in progressive …

Image Editing

Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation

2025-12-22 · Guoli Jia, Junyao Hu, Xinwei Long, Kai Tian 외 arxiv

Image generation based on diffusion models has demonstrated impressive capability, motivating exploration into diverse and specialized applications. Owing to the importance of emotion in advertising, emotion-oriented ima…

Image Generation