paper-with-me

홈 › Papers

MSGFusion: Multimodal Scene Graph-Guided Infrared and Visible Image Fusion

2025-09-16 · Guihui Li, Bowei Dong, Kaizhi Dong, Jiayi Li, Haiyong Zheng arxiv

Infrared and visible image fusion has garnered considerable attention owing to the strong complementarity of these two modalities in complex, harsh environments. While deep learning-based fusion methods have made remarkable advances in feature extraction, alignment, fusion, and reconstruction, they still depend largely on low-level visual cues, such as texture and contrast, and struggle to capture the high-level semantic information embedded in images. Recent attempts to incorporate text as a source of semantic guidance have relied on unstructured descriptions that neither explicitly model entities, attributes, and relationships nor provide spatial localization, thereby limiting fine-grained fusion performance. To overcome these challenges, we introduce MSGFusion, a multimodal scene graph-guided fusion framework for infrared and visible imagery. By deeply coupling structured scene graphs derived from text and vision, MSGFusion explicitly represents entities, attributes, and spatial relations, and then synchronously refines high-level semantics and low-level details through successive modules for scene graph representation, hierarchical aggregation, and graph-driven fusion. Extensive experiments on multiple public benchmarks show that MSGFusion significantly outperforms state-of-the-art approaches, particularly in detail preservation and structural clarity, and delivers superior semantic consistency and generalizability in downstream tasks such as low-light object detection, semantic segmentation, and medical image fusion.

📄 PDF Abstract BibTeX arXiv:2509.12901

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationObject Detection

Similar Papers 제목 키워드 기반

A Guided Upsampling Network for Short Wave Infrared Images Using Graph Regularization

2023-12-14 · Frank Sippel, Jürgen Seiler, André Kaup

Exploiting the infrared area of the spectrum for classification problems is getting increasingly popular, because many materials have characteristic absorption bands in this area. However, sensors in the short wave infra…

SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image Pretraining

2025-01-01 · CVPR 2025 1 · Mingjin Zhang, Xiaolong Li, Fei Gao, Jie Guo 외

Infrared Small Target Detection (IRSTD) aims to identify low signal-to-noise ratio small targets in infrared images with complex backgrounds, which is crucial for various applications. However, existing IRSTD methods…

Scene Recognition

Learning Spectral and Polarimetric Clues for One-to-Multimodal Novel View Synthesis

2026-07-02 · Federico Lincetto, Gianluca Agresti, Mattia Rossi, Piergiorgio Sartor 외 arxiv

Neural rendering techniques allow for accurate reconstruction of the geometry and color appearance of 3D scenes. Some methods have extended their use to additional imaging modalities, such as multispectral, infrared, or …

Novel View Synthesis

Fusing in 3D: Free-Viewpoint Fusion Rendering with a 3D Infrared-Visible Scene Representation

2026-01-19 · Chao Yang, Deshui Miao, Chao Tian, Guoqing Zhu 외 arxiv

Infrared-visible image fusion aims to integrate infrared and visible information into a single fused image. Existing 2D fusion methods focus on fusing images from fixed camera viewpoints, neglecting a comprehensive under…

Unified Restoration-Perception Learning: Maritime Infrared-Visible Image Fusion and Segmentation

2026-03-30 · Weichao Cai, Weiliang Huang, Biao Xue, Chao Huang 외 arxiv

Marine scene understanding and segmentation plays a vital role in maritime monitoring and navigation safety. However, prevalent factors like fog and strong reflections in maritime environments cause severe image degradat…

Semantic SegmentationScene UnderstandingImage Restoration