paper-with-me

홈 › Papers

Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching

2025-07-09 · Yafei Zhang, Yongle Shang, Huafeng Li arxiv

Weakly supervised text-to-person image matching, as a crucial approach to reducing models' reliance on large-scale manually labeled samples, holds significant research value. However, existing methods struggle to predict complex one-to-many identity relationships, severely limiting performance improvements. To address this challenge, we propose a local-and-global dual-granularity identity association mechanism. Specifically, at the local level, we explicitly establish cross-modal identity relationships within a batch, reinforcing identity constraints across different modalities and enabling the model to better capture subtle differences and correlations. At the global level, we construct a dynamic cross-modal identity association network with the visual modality as the anchor and introduce a confidence-based dynamic adjustment mechanism, effectively enhancing the model's ability to identify weakly associated samples while improving overall sensitivity. Additionally, we propose an information-asymmetric sample pair construction method combined with consistency learning to tackle hard sample mining and enhance model robustness. Experimental results demonstrate that the proposed method substantially boosts cross-modal matching accuracy, providing an efficient and practical solution for text-to-person image matching.

📄 PDF Abstract BibTeX arXiv:2507.06744

Code (0)

등록된 구현이 없습니다.

Tasks

Image Matching

Similar Papers 제목 키워드 기반

ViTag: Online WiFi Fine Time Measurements Aided Vision-Motion Identity Association in Multi-person Environments

2022-09-20 · IEEE International Conference on Sensing, Communication, and Networking 2022 9 · Bryan Bo Cao, Abrar Alali, Hansi Liu, Nicholas Meegan 외

In this paper, we present ViTag to associate user identities across multimodal data, particularly those obtained from cameras and smartphones. ViTag associates a sequence of vision tracker generated bounding boxes with I…

DecoderMultimodal AssociationTranslation

Audio-Visual Activity Guided Cross-Modal Identity Association for Active Speaker Detection

2022-12-01 · Rahul Sharma, Shrikanth Narayanan

Active speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality. The two primary sources of information to derive such a speech-face r…

Active Speaker DetectionAudio-Visual Active Speaker Detection

Causal Bootstrapped Alignment for Unsupervised Video-Based Visible-Infrared Person Re-Identification

2026-04-17 · Shuang Li, Jiaxu Leng, Changjiang Kuang, Mingpi Tan 외 arxiv

VVI-ReID is a critical technique for all-day surveillance, where temporal information provides additional cues beyond static images. However, existing approaches rely heavily on fully supervised learning with expensive c…

Person Re-Identification

MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models

2024-07-24 · Siwei Wu, Kang Zhu, Yu Bai, Yiming Liang 외

Given the remarkable success that large visual language models (LVLMs) have achieved in image perception tasks, the endeavor to make LVLMs perceive the world like humans is drawing increasing attention. Current multi-mod…

Language Modelling

Cross-Modality Earth Mover's Distance for Visible Thermal Person Re-Identification

2022-03-03 · Yongguo Ling, Zhun Zhong, Donglin Cao, Zhiming Luo 외

Visible thermal person re-identification (VT-ReID) suffers from the inter-modality discrepancy and intra-identity variations. Distribution alignment is a popular solution for VT-ReID, which, however, is usually restricte…

Person Re-Identification