paper-with-me

홈 › Papers

Infrared and visible Image Fusion with Language-driven Loss in CLIP Embedding Space

2024-02-26 · Yuhao Wang, Lingjuan Miao, Zhiqiang Zhou, Lei Zhang, Yajun Qiao

Infrared-visible image fusion (IVIF) has attracted much attention owing to the highly-complementary properties of the two image modalities. Due to the lack of ground-truth fused images, the fusion output of current deep-learning based methods heavily depends on the loss functions defined mathematically. As it is hard to well mathematically define the fused image without ground truth, the performance of existing fusion methods is limited. In this paper, we first propose to use natural language to express the objective of IVIF, which can avoid the explicit mathematical modeling of fusion output in current losses, and make full use of the advantage of language expression to improve the fusion performance. For this purpose, we present a comprehensive language-expressed fusion objective, and encode relevant texts into the multi-modal embedding space using CLIP. A language-driven fusion model is then constructed in the embedding space, by establishing the relationship among the embedded vectors to represent the fusion objective and input image modalities. Finally, a language-driven loss is derived to make the actual IVIF aligned with the embedded language-driven fusion model via supervised training. Experiments show that our method can obtain much better fusion results than existing techniques.

📄 PDF Abstract BibTeX arXiv:2402.16267

Code (1)

wyhlaowang/LDFusion 공식 구현 pytorch

Tasks

Infrared And Visible Image Fusion

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image Segmentation

2025-09-16 · Siju Ma, Changsiyu Gong, Xiaofeng Fan, Yong Ma 외 arxiv

Text-driven infrared and visible image fusion has gained attention for enabling natural language to guide the fusion process. However, existing methods lack a goal-aligned task to supervise and evaluate how effectively t…

Referring ExpressionImage Segmentation

DetFusion: A Detection-driven Infrared and Visible Image Fusion Network

2022-10-01 · ACMMM 2022 10 · Yiming Sun, Bing Cao, Pengfei Zhu, QinGhua Hu

Infrared and visible image fusion aims to utilize the complementary information between the two modalities to synthesize a new image containing richer information. Most existing works have focused on how to better fuse t…

Infrared And Visible Image FusionObjectobject-detectionObject Detection

SimpleFusion: A Simple Fusion Framework for Infrared and Visible Images

2024-06-27 · Ming Chen, Yuxuan Cheng, Xinwei He, Xinyue Wang 외

Integrating visible and infrared images into one high-quality image, also known as visible and infrared image fusion, is a challenging yet critical task for many downstream vision tasks. Most existing works utilize pretr…

HSFusion: A high-level vision task-driven infrared and visible image fusion network via semantic and geometric domain transformation

2024-07-14 · Chengjie Jiang, Xiaowen Liu, Bowen Zheng, Lu Bai 외

Infrared and visible image fusion has been developed from vision perception oriented fusion methods to strategies which both consider the vision perception and high-level vision task. However, the existing task-driven me…

Infrared And Visible Image FusionSemantic Segmentation

ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion

2026-07-26 · Guo Yurong, He Yufei, Li Yonghao, Chang Dongliang 외 arxiv

Controllable infrared-visible image fusion aims to integrate complementary thermal and structural information with flexible region-aware modulation, producing fused images that adapt to diverse user requirements and down…