paper-with-me

홈 › Papers

STA: Spatial-Temporal Attention for Large-Scale Video-based Person Re-Identification

2018-11-09 · Yang Fu, Xiaoyang Wang, Yunchao Wei, Thomas Huang

In this work, we propose a novel Spatial-Temporal Attention (STA) approach to tackle the large-scale person re-identification task in videos. Different from the most existing methods, which simply compute representations of video clips using frame-level aggregation (e.g. average pooling), the proposed STA adopts a more effective way for producing robust clip-level feature representation. Concretely, our STA fully exploits those discriminative parts of one target person in both spatial and temporal dimensions, which results in a 2-D attention score matrix via inter-frame regularization to measure the importances of spatial parts across different frames. Thus, a more robust clip-level feature representation can be generated according to a weighted sum operation guided by the mined 2-D attention score matrix. In this way, the challenging cases for video-based person re-identification such as pose variation and partial occlusion can be well tackled by the STA. We conduct extensive experiments on two large-scale benchmarks, i.e. MARS and DukeMTMC-VideoReID. In particular, the mAP reaches 87.7% on MARS, which significantly outperforms the state-of-the-arts with a large margin of more than 11.6%.

📄 PDF Abstract BibTeX arXiv:1811.04129

Code (0)

등록된 구현이 없습니다.

Tasks

Large-Scale Person Re-IdentificationPerson Re-IdentificationVideo-Based Person Re-Identification

Similar Papers 제목 키워드 기반

Where and when to look? Spatial-temporal attention for action recognition in videos

2019-05-01 · ICLR 2019 5 · Lili Meng, Bo Zhao, Bo Chang, Gao Huang 외

Inspired by the observation that humans are able to process videos efficiently by only paying attention when and where it is needed, we propose a novel spatial-temporal attention mechanism for video-based action recognit…

Action RecognitionAction Recognition In VideosTemporal Action LocalizationVideo Classification

TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video

2024-11-27 · Jinyuan Qu, Hongyang Li, Shilong Liu, Tianhe Ren 외

In this paper, we present TAPTRv3, which is built upon TAPTRv2 to improve its point tracking robustness in long videos. TAPTRv2 is a simple DETR-like framework that can accurately track any point in real-world videos wit…

Point Tracking

Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

2023-05-18 · Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He 외

With the explosive popularity of AI-generated content (AIGC), video generation has recently received a lot of attention. Generating videos guided by text instructions poses significant challenges, such as modeling the co…

Image GenerationText to Image GenerationText-to-Image GenerationText-to-Video Generation+2

R-STAN: Residual Spatial-Temporal Attention Network for Action Recognition

2019-06-19 · IEEE Access ( Volume: 7 ) 2019 6 · Quanle Liu, Xiangjiu Che, Mei Bie

Two-stream network architecture has the ability to capture temporal and spatial features from videos simultaneously and has achieved excellent performance on video action recognition tasks. However, there is a fair amoun…

Action RecognitionTemporal Action Localization

Video Crowd Localization with Multi-focus Gaussian Neighborhood Attention and a Large-Scale Benchmark

2021-07-19 · Haopeng Li, Lingbo Liu, Kunlin Yang, Shinan Liu 외

Video crowd localization is a crucial yet challenging task, which aims to estimate exact locations of human heads in the given crowded videos. To model spatial-temporal dependencies of human mobility, we propose a multi-…