Incremental Tube Construction for Human Action Detection
Current state-of-the-art action detection systems are tailored for offline batch-processing applications. However, for online applications like human-robot interaction, current systems fall short, either because they only detect one action per video, or because they assume that the entire video is available ahead of time. In this work, we introduce a real-time and online joint-labelling and association algorithm for action detection that can incrementally construct space-time action tubes on the most challenging action videos in which different action categories occur concurrently. In contrast to previous methods, we solve the detection-window association and action labelling problems jointly in a single pass. We demonstrate superior online association accuracy and speed (2.2ms per frame) as compared to the current state-of-the-art offline systems. We further demonstrate that the entire action detection pipeline can easily be made to work effectively in real-time using our action tube construction algorithm.
Code (1)
Tasks
Action DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Two-Stream AMTnet for Action Detection
In this paper, we propose Two-Stream AMTnet, which leverages recent advances in video-based action representation[1] and incremental action tube generation[2]. Majority of the present action detectors follow a frame-base…
Action DetectionAutonomous DrivingOnline Action DetectionOptical Flow Estimation+1TubeLite: Lightweight Multi-Actor Spatio-Temporal Action Detection
Spatio-temporal action detection in videos requires jointly localizing actors in space and identifying action boundaries over time. A common challenge is constructing temporally stable action tubes, as frame-level detect…
Action DetectionVideo-based Human-Object Interaction Detection from Tubelet Tokens
We present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatiotemporal representations, for video-based human-object interaction (V-HOI) detection. The tubel…
Human-Object Interaction DetectionDetecting Parts for Action Localization
In this paper, we propose a new framework for action localization that tracks people in videos and extracts full-body human tubes, i.e., spatio-temporal regions localizing actions, even in the case of occlusions or trunc…
Action LocalizationHuman Action Localization with Sparse Spatial Supervision
We introduce an approach for spatio-temporal human action localization using sparse spatial supervision. Our method leverages the large amount of annotated humans available today and extracts human tubes by combining a s…
Action LocalizationDiversity