Long Short-Term Transformer for Online Action Detection
We present Long Short-term TRansformer (LSTR), a temporal modeling algorithm for online action detection, which employs a long- and short-term memory mechanism to model prolonged sequence data. It consists of an LSTR encoder that dynamically leverages coarse-scale historical information from an extended temporal window (e.g., 2048 frames spanning of up to 8 minutes), together with an LSTR decoder that focuses on a short time window (e.g., 32 frames spanning 8 seconds) to model the fine-scale characteristics of the data. Compared to prior work, LSTR provides an effective and efficient method to model long videos with fewer heuristics, which is validated by extensive empirical analysis. LSTR achieves state-of-the-art performance on three standard online action detection benchmarks, THUMOS'14, TVSeries, and HACS Segment. Code has been made available at: https://xumingze0308.github.io/projects/lstr
Code (2)
Tasks
Action DetectionDecoderOnline Action DetectionPlaying the Game of 2048Similar Papers 제목 키워드 기반
Context-Enhanced Memory-Refined Transformer for Online Action Detection
Online Action Detection (OAD) detects actions in streaming videos using past observations. State-of-the-art OAD approaches model past observations and their interactions with an anticipated future. The past is encoded us…
Action DetectionDecoderOnline Action DetectionHAT: History-Augmented Anchor Transformer for Online Temporal Action Localization
Online video understanding often relies on individual frames, leading to frame-by-frame predictions. Recent advancements such as Online Temporal Action Localization (OnTAL), extend this approach to instance-level predict…
Action LocalizationTemporal Action LocalizationVideo UnderstandingReal-time Online Video Detection with Temporal Smoothing Transformers
Streaming video recognition reasons about objects and their actions in every frame of a video. A good streaming recognition model captures both long-term dynamics and short-term changes of video. Unfortunately, in most e…
Action AnticipationAction DetectionOnline Action DetectionVideo RecognitionFrom Implicit to Explicit feedback: A deep neural network for modeling sequential behaviours and long-short term preferences of online users
In this work, we examine the advantages of using multiple types of behaviour in recommendation systems. Intuitively, each user has to do some implicit actions (e.g., click) before making an explicit decision (e.g., purch…
Recommendation SystemsTransformer-based Nonlinear Equalization for DP-16QAM Coherent Optical Communication Systems
Compensating for nonlinear effects using digital signal processing (DSP) is complex and computationally expensive in long-haul optical communication systems due to intractable interactions between Kerr nonlinearity, chro…