Disentangle Object and Non-object Infrared Features via Language Guidance
Infrared object detection focuses on identifying and locating objects in complex environments (\eg, dark, snow, and rain) where visible imaging cameras are disabled by poor illumination. However, due to low contrast and weak edge information in infrared images, it is challenging to extract discriminative object features for robust detection. To deal with this issue, we propose a novel vision-language representation learning paradigm for infrared object detection. An additional textual supervision with rich semantic information is explored to guide the disentanglement of object and non-object features. Specifically, we propose a Semantic Feature Alignment (SFA) module to align the object features with the corresponding text features. Furthermore, we develop an Object Feature Disentanglement (OFD) module that disentangles text-aligned object features and non-object features by minimizing their correlation. Finally, the disentangled object features are entered into the detection head. In this manner, the detection performance can be remarkably enhanced via more discriminative and less noisy features. Extensive experimental results demonstrate that our approach achieves superior performance on two benchmarks: M\textsuperscript{3}FD (83.7\% mAP), FLIR (86.1\% mAP). Our code will be publicly available once the paper is accepted.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation LearningObject DetectionSimilar Papers 제목 키워드 기반
UIU-Net: U-Net in U-Net for Infrared Small Object Detection
Learning-based infrared small object detection methods currently rely heavily on the classification backbone network. This tends to result in tiny object loss and feature distinguishability limitations as the network dep…
Objectobject-detectionObject DetectionRepresentation Learning+1DPDETR: Decoupled Position Detection Transformer for Infrared-Visible Object Detection
Infrared-visible object detection aims to achieve robust object detection by leveraging the complementary information of infrared and visible image pairs. However, the commonly existing modality misalignment problem pres…
DecoderObjectobject-detectionObject Detection+2Deep Convolutional Neural Networks for Thermal Infrared Object Tracking
Unlike the visual object tracking, thermal infrared object tracking can track a target object in total darkness. Therefore, it has broad applications, such as in rescue and video surveillance at night. However, there are…
ObjectObject TrackingThermal Infrared Object TrackingVisual Object Tracking+1MetaFusion: Infrared and Visible Image Fusion via Meta-Feature Embedding From Object Detection
Fusing infrared and visible images can provide more texture details for subsequent object detection task. Conversely, detection task furnishes object semantic information to improve the infrared and visible image fus…
Infrared And Visible Image FusionMeta-LearningObjectobject-detection+1HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
Infrared small object detection is an important computer vision task involving the recognition and localization of tiny objects in infrared images, which usually contain only a few pixels. However, it encounters difficul…
channel selectionobject-detectionObject DetectionSmall Object Detection