SlowFast Rolling-Unrolling LSTMs for Action Anticipation in Egocentric Videos
Action anticipation in egocentric videos is a difficult task due to the inherently multi-modal nature of human actions. Additionally, some actions happen faster or slower than others depending on the actor or surrounding context which could vary each time and lead to different predictions. Based on this idea, we build upon RULSTM architecture, which is specifically designed for anticipating human actions, and propose a novel attention-based technique to evaluate, simultaneously, slow and fast features extracted from three different modalities, namely RGB, optical flow, and extracted objects. Two branches process information at different time scales, i.e., frame-rates, and several fusion schemes are considered to improve prediction accuracy. We perform extensive experiments on EpicKitchens-55 and EGTEA Gaze+ datasets, and demonstrate that our technique systematically improves the results of RULSTM architecture for Top-5 accuracy metric at different anticipation times.
Code (0)
등록된 구현이 없습니다.
Tasks
Action AnticipationOptical Flow EstimationRolling Shutter CorrectionSimilar Papers 제목 키워드 기반
Rolling-Unrolling LSTMs for Action Anticipation from First-Person Video
In this paper, we tackle the problem of egocentric action anticipation, i.e., predicting what actions the camera wearer will perform in the near future and which objects they will interact with. Specifically, we contribu…
Action AnticipationAction RecognitionOptical Flow EstimationRolling Shutter Correction+1What Would You Expect? Anticipating Egocentric Actions with Rolling-Unrolling LSTMs and Modality Attention
Egocentric action anticipation consists in understanding which objects the camera wearer will interact with in the near future and which actions they will perform. We tackle the problem proposing an architecture able to …
Action AnticipationAction RecognitionEgocentric Activity RecognitionOptical Flow Estimation+2Video + CLIP Baseline for Ego4D Long-term Action Anticipation
In this report, we introduce our adaptation of image-text models for long-term action anticipation. Our Video + CLIP framework makes use of a large-scale pre-trained paired image-text model: CLIP and a video encoder Slow…
Action AnticipationLong Term Action AnticipationTechnical Report for Ego4D Long Term Action Anticipation Challenge 2023
In this report, we describe the technical details of our approach for the Ego4D Long-Term Action Anticipation Challenge 2023. The aim of this task is to predict a sequence of future actions that will take place at an arb…
Action AnticipationDecoderLong Term Action AnticipationOperator Sketching for Deep Unrolling Networks
In this work we propose a new paradigm for designing efficient deep unrolling networks using operator sketching. The deep unrolling networks are currently the state-of-the-art solutions for imaging inverse problems. Howe…
Image ReconstructionRolling Shutter Correction