paper-with-me

Papers

Improving Description-based Person Re-identification by Multi-granularity Image-text Alignments

2019-06-23 · Kai Niu, Yan Huang, Wanli Ouyang, Liang Wang

Description-based person re-identification (Re-id) is an important task in video surveillance that requires discriminative cross-modal representations to distinguish different people. It is difficult to directly measure the similarity between images and descriptions due to the modality heterogeneity (the cross-modal problem). And all samples belonging to a single category (the fine-grained problem) makes this task even harder than the conventional image-description matching task. In this paper, we propose a Multi-granularity Image-text Alignments (MIA) model to alleviate the cross-modal fine-grained problem for better similarity evaluation in description-based person Re-id. Specifically, three different granularities, i.e., global-global, global-local and local-local alignments are carried out hierarchically. Firstly, the global-global alignment in the Global Contrast (GC) module is for matching the global contexts of images and descriptions. Secondly, the global-local alignment employs the potential relations between local components and global contexts to highlight the distinguishable components while eliminating the uninvolved ones adaptively in the Relation-guided Global-local Alignment (RGA) module. Thirdly, as for the local-local alignment, we match visual human parts with noun phrases in the Bi-directional Fine-grained Matching (BFM) module. The whole network combining multiple granularities can be end-to-end trained without complex pre-processing. To address the difficulties in training the combination of multiple granularities, an effective step training strategy is proposed to train these granularities step-by-step. Extensive experiments and analysis have shown that our method obtains the state-of-the-art performance on the CUHK-PEDES dataset and outperforms the previous methods by a significant margin.

📄 PDF Abstract BibTeX arXiv:1906.09610

Code (0)

등록된 구현이 없습니다.

Tasks

Image DescriptionPerson Re-IdentificationText based Person Retrieval

Similar Papers 제목 키워드 기반

Learning Granularity-Unified Representations for Text-to-Image Person Re-identification

2022-07-16 · Zhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin 외

Text-to-image person re-identification (ReID) aims to search for pedestrian images of an interested identity via textual descriptions. It is challenging due to both rich intra-modal variations and significant inter-modal…

Person Re-IdentificationText based Person RetrievalText based Person Search

Pose-Guided Multi-Granularity Attention Network for Text-Based Person Search

2018-09-22 · Ya Jing, Chenyang Si, Jun-Bo Wang, Wei Wang 외

Text-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surve…

Person SearchSentenceText based Person Search

Image-Text-Image Knowledge Transferring for Lifelong Person Re-Identification with Hybrid Clothing States

2024-05-26 · Qizao Wang, Xuelin Qian, Bin Li, Yanwei Fu 외

With the continuous expansion of intelligent surveillance networks, lifelong person re-identification (LReID) has received widespread attention, pursuing the need of self-evolution across different domains. However, exis…

Lifelong learningPerson Re-IdentificationTransfer Learning

Unsupervised Domain Adaptation for Cross-Regional Scenes Person Re-identification

2023-03-15 · Shanghai Jiao Tong Univ 2023 3 · Mao Yanmei, Li Huafeng, Zhang Yafei

In large-scale surveillance systems, the absence of positive cross-camera pedestrian samples in cross-regional scenes poses a limitation on the performance of person re-identification models. To tackle this challenge, an…

Domain AdaptationDomain Adaptive Person Re-IdentificationPerson Re-IdentificationStyle Transfer+1

Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images

2025-05-06 · Zengli Luo, Canlong Zhang, Xiaochun Lu, Zhixin Li 외

Text-based pedestrian search (TBPS) in full images aims to locate a target pedestrian in untrimmed images using natural language descriptions. However, in complex scenes with multiple pedestrians, existing methods are li…

Person SearchText based Person Search