Inductive Attention for Video Action Anticipation
Anticipating future actions based on spatiotemporal observations is essential in video understanding and predictive computer vision. Moreover, a model capable of anticipating the future has important applications, it can benefit precautionary systems to react before an event occurs. However, unlike in the action recognition task, future information is inaccessible at observation time -- a model cannot directly map the video frames to the target action to solve the anticipation task. Instead, the temporal inference is required to associate the relevant evidence with possible future actions. Consequently, existing solutions based on the action recognition models are only suboptimal. Recently, researchers proposed extending the observation window to capture longer pre-action profiles from past moments and leveraging attention to retrieve the subtle evidence to improve the anticipation predictions. However, existing attention designs typically use frame inputs as the query which is suboptimal, as a video frame only weakly connects to the future action. To this end, we propose an inductive attention model, dubbed IAM, which leverages the current prediction priors as the query to infer future action and can efficiently process the long video content. Furthermore, our method considers the uncertainty of the future via the many-to-many association in the attention design. As a result, IAM consistently outperforms the state-of-the-art anticipation models on multiple large-scale egocentric video datasets while using significantly fewer model parameters.
Code (0)
등록된 구현이 없습니다.
Tasks
Action AnticipationAction RecognitionVideo UnderstandingSimilar Papers 제목 키워드 기반
Knowledge-Guided Short-Context Action Anticipation in Human-Centric Videos
This work focuses on anticipating long-term human actions, particularly using short video segments, which can speed up editing workflows through improved suggestions while fostering creativity by suggesting narratives. T…
Action AnticipationLong Term Action AnticipationAction Anticipation from SoccerNet Football Video Broadcasts
Artificial intelligence has revolutionized the way we analyze sports videos, whether to understand the actions of games in long untrimmed videos or to anticipate the player's motion in future frames. Despite these effort…
Action AnticipationAction SpottingSports AnalyticsUntrimmed Action Anticipation
Egocentric action anticipation consists in predicting a future action the camera wearer will perform from egocentric video. While the task has recently attracted the attention of the research community, current approache…
Action AnticipationAction DetectionInteraction Region Visual Transformer for Egocentric Action Anticipation
Human-object interaction is one of the most important visual cues and we propose a novel way to represent human-object interactions for egocentric action anticipation. We propose a novel transformer variant to model inte…
Action AnticipationHuman-Object Interaction DetectionObjectFuture Transformer for Long-term Action Anticipation
The task of predicting future actions from a video is crucial for a real-world agent interacting with others. When anticipating actions in the distant future, we humans typically consider long-term relations over the who…
Action AnticipationLong Term Action AnticipationLong Term Anticipation