paper-with-me

홈 › Papers

Feature Disentanglement Learning with Switching and Aggregation for Video-based Person Re-Identification

2022-12-16 · Minjung Kim, MyeongAh Cho, Sangyoun Lee

In video person re-identification (Re-ID), the network must consistently extract features of the target person from successive frames. Existing methods tend to focus only on how to use temporal information, which often leads to networks being fooled by similar appearances and same backgrounds. In this paper, we propose a Disentanglement and Switching and Aggregation Network (DSANet), which segregates the features representing identity and features based on camera characteristics, and pays more attention to ID information. We also introduce an auxiliary task that utilizes a new pair of features created through switching and aggregation to increase the network's capability for various camera scenarios. Furthermore, we devise a Target Localization Module (TLM) that extracts robust features against a change in the position of the target according to the frame flow and a Frame Weight Generation (FWG) that reflects temporal information in the final representation. Various loss functions for disentanglement learning are designed so that each component of the network can cooperate while satisfactorily performing its own role. Quantitative and qualitative results from extensive experiments demonstrate the superiority of DSANet over state-of-the-art methods on three benchmark datasets.

📄 PDF Abstract BibTeX arXiv:2212.09498

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementPerson Re-IdentificationVideo-Based Person Re-Identification

Similar Papers 제목 키워드 기반

Frame Aggregation and Multi-Modal Fusion Framework for Video-Based Person Recognition

2020-10-19 · Fangtao Li, Wenzhe Wang, Zihe Liu, Haoran Wang 외

Video-based person recognition is challenging due to persons being blocked and blurred, and the variation of shooting angle. Previous research always focused on person recognition on still images, ignoring similarity and…

Person Recognition

Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-based Person Re-identification

2020-03-27 · CVPR 2020 6 · Zhizheng Zhang, Cuiling Lan, Wen-Jun Zeng, Zhibo Chen

Video-based person re-identification (reID) aims at matching the same person across video clips. It is a challenging task due to the existence of redundancy among frames, newly revealed appearance, occlusion, and motion …

Person Re-IdentificationVideo-Based Person Re-Identification

Shuffle Transformer with Feature Alignment for Video Face Parsing

2021-06-16 · Rui Zhang, Yang Han, Zilong Huang, Pei Cheng 외

This is a short technical report introducing the solution of the Team TCParser for Short-video Face Parsing Track of The 3rd Person in Context (PIC) Workshop and Challenge at CVPR 2021. In this paper, we introduce a stro…

Face Parsing

An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement

2024-02-27 · Tzu-Ting Yang, Hsin-Wei Wang, Yi-Cheng Wang, Chi-Han Lin 외

With the massive developments of end-to-end (E2E) neural networks, recent years have witnessed unprecedented breakthroughs in automatic speech recognition (ASR). However, the codeswitching phenomenon remains a major obst…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DisentanglementMixture-of-Experts+2

A Unified Latent Space Disentanglement VAE Framework with Robust Disentanglement Effectiveness Evaluation

2026-03-11 · Xiaoan Lang, Md Mostafizer Rahman, Fang Liu arxiv

Evaluating and interpreting latent representations, such as variational autoencoders (VAEs), remains a significant challenge for diverse data types, especially when ground-truth generative factors are unknown. To address…