Learning Semantic-Aligned Feature Representation for Text-based Person Search
Text-based person search aims to retrieve images of a certain pedestrian by a textual description. The key challenge of this task is to eliminate the inter-modality gap and achieve the feature alignment across modalities. In this paper, we propose a semantic-aligned embedding method for text-based person search, in which the feature alignment across modalities is achieved by automatically learning the semantic-aligned visual features and textual features. First, we introduce two Transformer-based backbones to encode robust feature representations of the images and texts. Second, we design a semantic-aligned feature aggregation network to adaptively select and aggregate features with the same semantics into part-aware features, which is achieved by a multi-head attention module constrained by a cross-modality part alignment loss and a diversity loss. Experimental results on the CUHK-PEDES and Flickr30K datasets show that our method achieves state-of-the-art performances.
Code (1)
Tasks
DiversityPerson SearchText based Person RetrievalText based Person SearchMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Semantics-Aligned Representation Learning for Person Re-identification
Person re-identification (reID) aims to match person images to retrieve the ones with the same identity. This is a challenging task, as the images to be matched are generally semantically misaligned due to the diversity …
DecoderPerson Re-IdentificationRepresentation LearningTexture Synthesis+1AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identification
Cross-modal person re-identification (Re-ID) is critical for modern video surveillance systems. The key challenge is to align cross-modality representations induced by the semantic information present for a person and ig…
Cross-Modal Person Re-IdentificationCross-Modal Person Re-IdentificationPerson Re-IdentificationPerson Search+2VGSG: Vision-Guided Semantic-Group Network for Text-based Person Search
Text-based Person Search (TBPS) aims to retrieve images of target pedestrian indicated by textual descriptions. It is essential for TBPS to extract fine-grained local features and align them crossing modality. Existing m…
Person SearchText based Person RetrievalText based Person SearchTransfer LearningSemantically Self-Aligned Network for Text-to-Image Part-aware Person Re-identification
Text-to-image person re-identification (ReID) aims to search for images containing a person of interest using textual descriptions. However, due to the significant modality gap and the large intra-class variance in textu…
Image RetrievalPerson Re-IdentificationText based Person RetrievalText-based Person Retrieval with Noisy CorrespondenceCo-Attention Aligned Mutual Cross-Attention for Cloth-Changing Person Re-Identification
Person re-identification (Re-ID) has been widely studied and achieved significant progress. However, traditional person Re-ID methods primarily rely on cloth-related color appearance, which is unreliable under real-world…
Cloth-Changing Person Re-IdentificationPerson Re-IdentificationPerson RetrievalRetrieval