paper-with-me

홈 › Papers

A Cross-Modal Image Fusion Method Guided by Human Visual Characteristics

2019-12-18 · Aiqing Fang, Xinbo Zhao, Jiaqi Yang, Yanning Zhang

The characteristics of feature selection, nonlinear combination and multi-task auxiliary learning mechanism of the human visual perception system play an important role in real-world scenarios, but the research of image fusion theory based on the characteristics of human visual perception is less. Inspired by the characteristics of human visual perception, we propose a robust multi-task auxiliary learning optimization image fusion theory. Firstly, we combine channel attention model with nonlinear convolutional neural network to select features and fuse nonlinear features. Then, we analyze the impact of the existing image fusion loss on the image fusion quality, and establish the multi-loss function model of unsupervised learning network. Secondly, aiming at the multi-task auxiliary learning mechanism of human visual perception system, we study the influence of multi-task auxiliary learning mechanism on image fusion task on the basis of single task multi-loss network model. By simulating the three characteristics of human visual perception system, the fused image is more consistent with the mechanism of human brain image fusion. Finally, in order to verify the superiority of our algorithm, we carried out experiments on the combined vision system image data set, and extended our algorithm to the infrared and visible image and the multi-focus image public data set for experimental verification. The experimental results demonstrate the superiority of our fusion theory over state-of-arts in generality and robustness.

📄 PDF Abstract BibTeX arXiv:1912.08577

Code (0)

등록된 구현이 없습니다.

Tasks

Auxiliary Learningfeature selection

Similar Papers 제목 키워드 기반

DiffX: Guide Your Layout to Cross-Modal Generative Modeling

2024-07-22 · Zeyu Wang, Jingyu Lin, Yifei Qian, Yi Huang 외

Diffusion models have made significant strides in language-driven and layout-driven image generation. However, most diffusion models are limited to visible RGB image generation. In fact, human perception of the world is …

DenoisingImage CaptioningImage Generation

PMMD: A pose-guided multi-view multi-modal diffusion for person generation

2025-12-17 · Ziyu Shang, Haoran Liu, Rongchao Zhang, Zhiqian Wei 외 arxiv

Generating consistent human images with controllable pose and appearance is essential for applications in virtual try on, image editing, and digital human creation. Current methods often suffer from occlusions, garment s…

Image Editing

Cross-Modal Image Fusion Theory Guided by Subjective Visual Attention

2019-12-23 · Aiqing Fang, Xinbo Zhao, Yanning Zhang

The human visual perception system has very strong robustness and contextual awareness in a variety of image processing tasks. This robustness and the perception ability of contextual awareness is closely related to the …

Auxiliary Learning

HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation

2022-11-11 · Kaiduo Zhang, Muyi Sun, Jianxin Sun, Binghao Zhao 외

Text-driven person image generation is an emerging and challenging task in cross-modality image generation. Controllable person image generation promotes a wide range of applications such as digital human interaction and…

Image GenerationRetrievalSentenceVirtual Try-on

Coupled Degradation Modeling and Fusion: A VLM-Guided Degradation-Coupled Network for Degradation-Aware Infrared and Visible Image Fusion

2025-10-13 · Tianpei Zhang, Jufeng Zhao, Yiming Zhu, Guangmang Cui arxiv

Existing Infrared and Visible Image Fusion (IVIF) methods typically assume high-quality inputs. However, when handing degraded images, these methods heavily rely on manually switching between different pre-processing tec…