Learning Person Re-identification Models from Videos with Weak Supervision
Most person re-identification methods, being supervised techniques, suffer from the burden of massive annotation requirement. Unsupervised methods overcome this need for labeled data, but perform poorly compared to the supervised alternatives. In order to cope with this issue, we introduce the problem of learning person re-identification models from videos with weak supervision. The weak nature of the supervision arises from the requirement of video-level labels, i.e. person identities who appear in the video, in contrast to the more precise framelevel annotations. Towards this goal, we propose a multiple instance attention learning framework for person re-identification using such video-level labels. Specifically, we first cast the video person re-identification task into a multiple instance learning setting, in which person images in a video are collected into a bag. The relations between videos with similar labels can be utilized to identify persons, on top of that, we introduce a co-person attention mechanism which mines the similarity correlations between videos with person identities in common. The attention weights are obtained based on all person images instead of person tracklets in a video, making our learned model less affected by noisy annotations. Extensive experiments demonstrate the superiority of the proposed method over the related methods on two weakly labeled person re-identification datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Multiple Instance LearningPerson Re-IdentificationVideo-Based Person Re-IdentificationSimilar Papers 제목 키워드 기반
Large-Scale Pre-training for Person Re-identification with Noisy Labels
This paper aims to address the problem of pre-training for person re-identification (Re-ID) with noisy labels. To setup the pre-training task, we apply a simple online multi-object tracking system on raw videos of an exi…
Contrastive LearningMulti-Object TrackingObject TrackingOnline Multi-Object Tracking+2Actor and Observer: Joint Modeling of First and Third-Person Videos
Several theories in cognitive neuroscience suggest that when people interact with the world, or simulate interactions, they do so from a first-person egocentric perspective, and seamlessly transfer knowledge between thir…
Action RecognitionTemporal Action LocalizationTransferring a Semantic Representation for Person Re-Identification and Search
Learning semantic attributes for person re-identification and description-based person search has gained increasing interest due to attributes' great potential as a pose and view-invariant representation. However, existi…
AttributePerson Re-IdentificationPerson SearchWeakly supervised discriminative feature learning with state information for person identification
Unsupervised learning of identity-discriminative visual feature is appealing in real-world tasks where manual labelling is costly. However, the images of an identity can be visually discrepant when images are taken under…
Face RecognitionPerson IdentificationPerson Re-IdentificationPseudo Label+2Weakly-Supervised Multi-Person Action Recognition in 360$^{\circ}$ Videos
The recent development of commodity 360$^{\circ}$ cameras have enabled a single video to capture an entire scene, which endows promising potentials in surveillance scenarios. However, research in omnidirectional video an…
Action LocalizationAction RecognitionMulti-Label Learning