Anticipative Feature Fusion Transformer for Multi-Modal Action Anticipation
Although human action anticipation is a task which is inherently multi-modal, state-of-the-art methods on well known action anticipation datasets leverage this data by applying ensemble methods and averaging scores of unimodal anticipation networks. In this work we introduce transformer based modality fusion techniques, which unify multi-modal data at an early stage. Our Anticipative Feature Fusion Transformer (AFFT) proves to be superior to popular score fusion approaches and presents state-of-the-art results outperforming previous methods on EpicKitchens-100 and EGTEA Gaze+. Our model is easily extensible and allows for adding new modalities without architectural changes. Consequently, we extracted audio features on EpicKitchens-100 which we add to the set of commonly used features in the community.
Code (1)
Tasks
Action AnticipationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Solar Irradiance Anticipative Transformer
This paper proposes an anticipative transformer-based model for short-term solar irradiance forecasting. Given a sequence of sky images, our proposed vision transformer encodes features of consecutive images, feeding int…
DecoderSolar Irradiance ForecastingOn the Efficacy of Text-Based Input Modalities for Action Anticipation
Anticipating future actions is a highly challenging task due to the diversity and scale of potential future actions; yet, information from different modalities help narrow down plausible action choices. Each modality can…
Action AnticipationAnticipative Video Transformer
We propose Anticipative Video Transformer (AVT), an end-to-end attention-based video modeling architecture that attends to the previously observed video in order to anticipate future actions. We train the model jointly t…
Action AnticipationAbout the decomposition of pricing formulas under stochastic volatility models
We obtain a decomposition of the call option price for a very general stochastic volatility diffusion model extending the decomposition obtained by E. Al\`os in [2] for the Heston model. We realize that a new term arises…
Text-Derived Knowledge Helps Vision: A Simple Cross-modal Distillation for Video-based Action Anticipation
Anticipating future actions in a video is useful for many autonomous and assistive technologies. Most prior action anticipation work treat this as a vision modality problem, where the models learn the task information pr…
Action AnticipationTransfer Learning