Keypoint-Aligned Embeddings for Image Retrieval and Re-identification
Learning embeddings that are invariant to the pose of the object is crucial in visual image retrieval and re-identification. The existing approaches for person, vehicle, or animal re-identification tasks suffer from high intra-class variance due to deformable shapes and different camera viewpoints. To overcome this limitation, we propose to align the image embedding with a predefined order of the keypoints. The proposed keypoint aligned embeddings model (KAE-Net) learns part-level features via multi-task learning which is guided by keypoint locations. More specifically, KAE-Net extracts channels from a feature map activated by a specific keypoint through learning the auxiliary task of heatmap reconstruction for this keypoint. The KAE-Net is compact, generic and conceptually simple. It achieves state of the art performance on the benchmark datasets of CUB-200-2011, Cars196 and VeRi-776 for retrieval and re-identification tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
Image RetrievalMulti-Task LearningRetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Skeleton Merger: an Unsupervised Aligned Keypoint Detector
Detecting aligned 3D keypoints is essential under many scenarios such as object tracking, shape retrieval and robotics. However, it is generally hard to prepare a high-quality dataset for all types of objects due to the …
DecoderObject TrackingRetrievalUnsupervised Keypoints from Pretrained Diffusion Models
Unsupervised learning of keypoints and landmarks has seen significant progress with the help of modern neural network architectures, but performance is yet to match the supervised counterpart, making their practicability…
DenoisingUnsupervised Human Pose EstimationUnsupervised KeypointsMEDIAPI-SKEL - A 2D-Skeleton Video Database of French Sign Language With Aligned French Subtitles
This paper presents MEDIAPI-SKEL, a 2D-skeleton database of French Sign Language videos aligned with French subtitles. The corpus contains 27 hours of video of body, face and hand keypoints, aligned to subtitles with a v…
Cross-Modal RetrievalRetrievalSemantic SegmentationVideo Semantic SegmentationPentagon-Match (PMatch): Identification of View-Invariant Planar Feature for Local Feature Matching-Based Homography Estimation
In computer vision, finding correct point correspondence among images plays an important role in many applications, such as image stitching, image retrieval, visual localization, etc. Most of the research works focus on …
Homography EstimationImage RetrievalImage StitchingRetrieval+1Keypoint Promptable Re-Identification
Occluded Person Re-Identification (ReID) is a metric learning task that involves matching occluded individuals based on their appearance. While many studies have tackled occlusions caused by objects, multi-person occlusi…
Metric LearningOccluded Person Re-IdentificationPerson Re-IdentificationPerson Retrieval+1