paper-with-me

홈 › Papers

Distilling Textual Priors from LLM to Efficient Image Fusion

2025-04-09 · Ran Zhang, Xuanhua He, Ke Cao, Liu Liu, Li Zhang, Man Zhou, Jie Zhang

Multi-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but struggle to handle low-quality or complex inputs. Recent advances in text-guided methods leverage large model priors to overcome these limitations, but at the cost of significant computational overhead, both in memory and inference time. To address this challenge, we propose a novel framework for distilling large model priors, eliminating the need for text guidance during inference while dramatically reducing model size. Our framework utilizes a teacher-student architecture, where the teacher network incorporates large model priors and transfers this knowledge to a smaller student network via a tailored distillation process. Additionally, we introduce spatial-channel cross-fusion module to enhance the model's ability to leverage textual priors across both spatial and channel dimensions. Our method achieves a favorable trade-off between computational efficiency and fusion quality. The distilled network, requiring only 10% of the parameters and inference time of the teacher network, retains 90% of its performance and outperforms existing SOTA methods. Extensive experiments demonstrate the effectiveness of our approach. The implementation will be made publicly available as an open-source resource.

📄 PDF Abstract BibTeX arXiv:2504.07029

Code (1)

zirconium233/dtpf 공식 구현 pytorch

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse Views

2023-08-27 · Zi-Xin Zou, Weihao Cheng, Yan-Pei Cao, Shi-Sheng Huang 외

Reconstructing 3D objects from extremely sparse views is a long-standing and challenging problem. While recent techniques employ image diffusion models for generating plausible images at novel viewpoints or for distillin…

3D ReconstructionNovel View SynthesisObject Reconstruction

GeoDream: Disentangling 2D and Geometric Priors for High-Fidelity and Consistent 3D Generation

2023-11-29 · Baorui Ma, Haoge Deng, Junsheng Zhou, Yu-Shen Liu 외

Text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models has shown great promise but still suffers from inconsistent 3D geometric structures (Janus problems) and severe artifacts. The afo…

3D GenerationText to 3D

HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement

2026-04-12 · Marco Schouten, Ioannis Siglidis, Serge Belongie, Dim P. Papadopoulos arxiv

We propose a method to learn explicit, class-conditioned spatial priors for object placement in natural scenes by distilling the implicit placement knowledge encoded in text-conditioned diffusion models. Prior work relie…

Image Editing

Distilling Semantic Priors from SAM to Efficient Image Restoration Models

2024-03-25 · CVPR 2024 1 · Quan Zhang, Xiaoyu Liu, Wei Li, Hanting Chen 외

In image restoration (IR), leveraging semantic priors from segmentation models has been a common approach to improve performance. The recent segment anything model (SAM) has emerged as a powerful tool for extracting adva…

DeblurringDenoisingImage RestorationRain Removal

GaussianSR: 3D Gaussian Super-Resolution with 2D Diffusion Priors

2024-06-14 · Xiqian Yu, Hanxin Zhu, Tianyu He, Zhibo Chen

Achieving high-resolution novel view synthesis (HRNVS) from low-resolution input views is a challenging task due to the lack of high-resolution data. Previous methods optimize high-resolution Neural Radiance Field (NeRF)…

3DGSNeRFNovel View SynthesisSuper-Resolution