paper-with-me

홈 › Papers

DreamFuse: Adaptive Image Fusion with Diffusion Transformer

2025-04-11 · Junjia Huang, Pengxiang Yan, Jiyang Liu, Jie Wu, Zhao Wang, Yitong Wang, Liang Lin, Guanbin Li

Image fusion seeks to seamlessly integrate foreground objects with background scenes, producing realistic and harmonious fused images. Unlike existing methods that directly insert objects into the background, adaptive and interactive fusion remains a challenging yet appealing task. It requires the foreground to adjust or interact with the background context, enabling more coherent integration. To address this, we propose an iterative human-in-the-loop data generation pipeline, which leverages limited initial data with diverse textual prompts to generate fusion datasets across various scenarios and interactions, including placement, holding, wearing, and style transfer. Building on this, we introduce DreamFuse, a novel approach based on the Diffusion Transformer (DiT) model, to generate consistent and harmonious fused images with both foreground and background information. DreamFuse employs a Positional Affine mechanism to inject the size and position of the foreground into the background, enabling effective foreground-background interaction through shared attention. Furthermore, we apply Localized Direct Preference Optimization guided by human feedback to refine DreamFuse, enhancing background consistency and foreground harmony. DreamFuse achieves harmonious fusion while generalizing to text-driven attribute editing of the fused results. Experimental results demonstrate that our method outperforms state-of-the-art approaches across multiple metrics.

📄 PDF Abstract BibTeX arXiv:2504.08291

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeStyle Transfer

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On

2025-03-26 · CVPR 2025 1 · Ji Woo Hong, Tri Ton, Trung X. Pham, Gwanhyeong Koo 외

This paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of previous approaches by leveraging the Mask…

DenoisingVirtual Try-on

Effective Diffusion Transformer Architecture for Image Super-Resolution

2024-09-29 · Kun Cheng, Lei Yu, Zhijun Tu, Xiao He 외

Recent advances indicate that diffusion models hold great promise in image super-resolution. While the latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attem…

Image GenerationImage Super-ResolutionSuper-Resolution

AdaCorrection: Adaptive Offset Cache Correction for Accurate Diffusion Transformers

2026-02-13 · Dong Liu, Yanxuan Yu, Ben Lengerich, Ying Nian Wu arxiv

Diffusion Transformers (DiTs) achieve state-of-the-art performance in high-fidelity image and video generation but suffer from expensive inference due to their iterative denoising structure. While prior methods accelerat…

Video Generation

EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing

2024-10-02 · Haotian Sun, Tao Lei, BoWen Zhang, Yanghao Li 외

Diffusion transformers have been widely adopted for text-to-image synthesis. While scaling these models up to billions of parameters shows promise, the effectiveness of scaling beyond current sizes remains underexplored …

Image GenerationMixture-of-Experts

Quality-Aware Modulation for Diffusion Transformers

2026-06-29 · Luke Budny, Yuhong Guo, Kevin Cheung arxiv

Modern text-to-image diffusion models, such as diffusion transformers (DiT), rely on timestep or prompt embeddings to modulate the strength of the denoising process in each timestep. While this modulation communicates th…