Attribute-Driven Feature Disentangling and Temporal Aggregation for Video Person Re-Identification
Video-based person re-identification plays an important role in surveillance video analysis, expanding image-based methods by learning features of multiple frames. Most existing methods fuse features by temporal average-pooling, without exploring the different frame weights caused by various viewpoints, poses, and occlusions. In this paper, we propose an attribute-driven method for feature disentangling and frame re-weighting. The features of single frames are disentangled into groups of sub-features, each corresponds to specific semantic attributes. The sub-features are re-weighted by the confidence of attribute recognition and then aggregated at the temporal dimension as the final representation. By means of this strategy, the most informative regions of each frame are enhanced and contributes to a more discriminative sequence representation. Extensive ablation studies demonstrate the effectiveness of feature disentangling as well as temporal re-weighting. The experimental results on the iLIDS-VID, PRID-2011 and MARS datasets demonstrate that our proposed method outperforms existing state-of-the-art approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributePerson Re-IdentificationVideo-Based Person Re-IdentificationSimilar Papers 제목 키워드 기반
Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition
Spatial-temporal graphs have been widely used by skeleton-based action recognition algorithms to model human action dynamics. To capture robust movement patterns from these graphs, long-range and multi-scale context aggr…
3D Action RecognitionAction RecognitionLong-range modelingSkeleton Based Action RecognitionAttribute-driven Disentangled Representation Learning for Multimodal Recommendation
Recommendation algorithms forecast user preferences by correlating user and item representations derived from historical interaction patterns. In pursuit of enhanced performance, many methods focus on learning robust and…
AttributeMultimodal RecommendationRepresentation LearningMulti-Stage Spatio-Temporal Aggregation Transformer for Video Person Re-identification
In recent years, the Transformer architecture has shown its superiority in the video-based person re-identification task. Inspired by video representation learning, these methods mainly focus on designing modules to extr…
AttributePerson Re-IdentificationRepresentation LearningVideo-Based Person Re-IdentificationAttribute-Based Progressive Fusion Network for RGBT Tracking
RGBT tracking usually suffers from various challenging factors of fast motion, scale variation, illumination variation,thermal crossover and occlusion, to name a few. Existing works often study fusion models to solve all…
AttributeRgb-T TrackingFTM: A Frame-level Timeline Modeling Method for Temporal Graph Representation Learning
Learning representations for graph-structured data is essential for graph analytical tasks. While remarkable progress has been made on static graphs, researches on temporal graphs are still in its beginning stage. The bo…
Graph Representation LearningRepresentation Learning