paper-with-me

Papers

CLIP-Based Modality Compensation for Visible-Infrared Image Re-Identification

2024-12-25 · journal 2024 12 · Gang Hu; Yafei Lv; Jianting Zhang; Qian Wu; Zaidao Wen

Visible-infrared image re-identification (VIReID) aims to match objects with the same identity appearing across different modalities. Given the significant differences between visible and infrared images, VIReID poses a formidable challenge. Most existing methods focus on extracting modality-shared features while ignore modality-specific features, which often also contain crucial important discriminative information. In addition, high-level semantic information of the objects, such as shape and appearance, is also crucial for the VIReID task. To further enhance the retrieval performance, we propose a novel one-stage CLIP-based Modality Compensation (CLIP-MC) method for the VIReID task. Our method introduces a new prompt learning paradigm that leverages the semantic understanding capabilities of CLIP to recover missing modality information. CLIP-MC comprises three key modules: Instance Text Prompt Generation (ITPG), Modality Compensation (MC), and Modality Context Learner (MCL). Specifically, the ITPG module facilitates effective alignment and interaction between image tokens and text tokens, enhancing the text encoder's ability to capture detailed visual information from the images. This ensures that the text encoder generates fine-grained descriptions of the images. The MCL module captures the unique information of each modality and generates modality-specific context tokens, which are more flexible compared to fixed text descriptions. Guided by the modality-specific context, the text encoder discovers missing modality information from the images and produces compensated modality features. Finally, the MC module combines the original and compensated modality features to obtain complete modality features that contain more discriminative information. We conduct extensive experiments on three VIReID datasets and compare the performance of our method with other existing approaches to demonstrate its effectiveness and superiority.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal Person Re-IdentificationCross-Modal Person Re-IdentificationPerson Re-IdentificationPrompt Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Focus 설명 없음

Similar Papers 제목 키워드 기반

FMCNet: Feature-Level Modality Compensation for Visible-Infrared Person Re-Identification

2022-01-01 · CVPR 2022 1 · Qiang Zhang, Changzhou Lai, Jianan Liu, Nianchang Huang 외

For Visible-Infrared person Re-IDentification (VI-ReID), existing modality-specific information compensation based models try to generate the images of missing modality from existing ones for reducing cross-modality …

Person Re-Identification

CLIP4VI-ReID: Learning Modality-shared Representations via CLIP Semantic Bridge for Visible-Infrared Person Re-identification

2025-11-13 · Xiaomei Yang, Xizhan Gao, Sijie Niu, Fa Zhu 외 arxiv

This paper proposes a novel CLIP-driven modality-shared representation learning network named CLIP4VI-ReID for VI-ReID task, which consists of Text Semantic Generation (TSG), Infrared Feature Embedding (IFE), and High-le…

Person Re-IdentificationRepresentation Learning

MRCN: A Novel Modality Restitution and Compensation Network for Visible-Infrared Person Re-identification

2023-03-26 · Yukang Zhang, Yan Yan, Jie Li, Hanzi Wang

Visible-infrared person re-identification (VI-ReID), which aims to search identities across different spectra, is a challenging task due to large cross-modality discrepancy between visible and infrared images. The key to…

Person Re-Identification

CLIP-Driven Semantic Discovery Network for Visible-Infrared Person Re-Identification

2024-01-11 · Xiaoyan Yu, Neng Dong, Liehuang Zhu, Hao Peng 외

Visible-infrared person re-identification (VIReID) primarily deals with matching identities across person images from different modalities. Due to the modality gap between visible and infrared images, cross-modality iden…

Person Re-Identification

MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation

2026-07-07 · Jiaju Han, Ma Yaqi, Yahui Chai, Xuemeng Sun 외 arxiv

Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focu…