paper-with-me

홈 › Papers

Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach

2025-12-08 · Jiayang Li, Chengjie Jiang, Junjun Jiang, Pengwei Liang, Jiayi Ma, Liqiang Nie arxiv

Image fusion aims to blend complementary information from multiple sensing modalities, yet existing approaches remain limited in robustness, adaptability, and controllability. Most current fusion networks are tailored to specific tasks and lack the ability to flexibly incorporate user intent, especially in complex scenarios involving low-light degradation, color shifts, or exposure imbalance. Moreover, the absence of ground-truth fused images and the small scale of existing datasets make it difficult to train an end-to-end model that simultaneously understands high-level semantics and performs fine-grained multimodal alignment. We therefore present DiTFuse, instruction-driven Diffusion-Transformer (DiT) framework that performs end-to-end, semantics-aware fusion within a single model. By jointly encoding two images and natural-language instructions in a shared latent space, DiTFuse enables hierarchical and fine-grained control over fusion dynamics, overcoming the limitations of pre-fusion and post-fusion pipelines that struggle to inject high-level semantics. The training phase employs a multi-degradation masked-image modeling strategy, so the network jointly learns cross-modal alignment, modality-invariant restoration, and task-aware feature selection without relying on ground truth images. A curated, multi-granularity instruction dataset further equips the model with interactive fusion capabilities. DiTFuse unifies infrared-visible, multi-focus, and multi-exposure fusion-as well as text-controlled refinement and downstream tasks-within a single architecture. Experiments on public IVIF, MFF, and MEF benchmarks confirm superior quantitative and qualitative performance, sharper textures, and better semantic retention. The model also supports multi-level user control and zero-shot generalization to other multi-image fusion scenarios, including instruction-conditioned segmentation.

📄 PDF Abstract BibTeX arXiv:2512.07170

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot Generalization

Similar Papers 제목 키워드 기반

More Control for Free! Image Synthesis with Semantic Diffusion Guidance

2021-12-10 · Xihui Liu, Dong Huk Park, Samaneh Azadi, Gong Zhang 외

Controllable image synthesis models allow creation of diverse images based on text instructions or guidance from a reference image. Recently, denoising diffusion probabilistic models have been shown to generate more real…

continuous-controlContinuous ControlDenoisingImage Generation

A2BFR: Attribute-Aware Blind Face Restoration

2026-03-31 · Chenxin Zhu, Yushun Fang, Lu Liu, Shibo Yin 외 arxiv

Blind face restoration (BFR) aims to recover high-quality facial images from degraded inputs, yet its inherently ill-posed nature leads to ambiguous and uncontrollable solutions. Recent diffusion-based BFR methods improv…

Blind Face Restoration

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

2026-06-29 · Qin Guo, Hao Luo, Dongxu Yue, Weixuan Jin 외 arxiv

Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches typically treat diffusion-based controllable generation and …

Image Generation

MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation

2026-03-30 · Bharath Krishnamurthy, Ajita Rattani arxiv

Recent multimodal face generation models address the spatial control limitations of text-to-image diffusion models by augmenting text-based conditioning with spatial priors such as segmentation masks, sketches, or edge m…

Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion

2026-04-27 · Zhongjie Duan, Hong Zhang, Yingda Chen arxiv

Controllable diffusion methods have substantially expanded the practical utility of diffusion models, but they are typically developed as isolated, backbone-specific systems with incompatible training pipelines, paramete…

Image Editing