paper-with-me

홈 › Papers

On the Efficacy of Text-Based Input Modalities for Action Anticipation

2024-01-23 · Apoorva Beedu, Harish Haresamudram, Karan Samel, Irfan Essa

Anticipating future actions is a highly challenging task due to the diversity and scale of potential future actions; yet, information from different modalities help narrow down plausible action choices. Each modality can provide diverse and often complementary context for the model to learn from. While previous multi-modal methods leverage information from modalities such as video and audio, we primarily explore how text descriptions of actions and objects can also lead to more accurate action anticipation by providing additional contextual cues, e.g., about the environment and its contents. We propose a Multi-modal Contrastive Anticipative Transformer (M-CAT), a video transformer architecture that jointly learns from multi-modal features and text descriptions of actions and objects. We train our model in two stages, where the model first learns to align video clips with descriptions of future actions, and is subsequently fine-tuned to predict future actions. Compared to existing methods, M-CAT has the advantage of learning additional context from two types of text inputs: rich descriptions of future actions during pre-training, and, text descriptions for detected objects and actions during modality feature fusion. Through extensive experimental evaluation, we demonstrate that our model outperforms previous methods on the EpicKitchens datasets, and show that using simple text descriptions of actions and objects aid in more effective action anticipation. In addition, we examine the impact of object and action information obtained via text, and perform extensive ablations.

📄 PDF Abstract BibTeX arXiv:2401.12972

Code (0)

등록된 구현이 없습니다.

Tasks

Action Anticipation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Attention 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.

Similar Papers 제목 키워드 기반

Cross-modal Contrastive Distillation for Instructional Activity Anticipation

2022-01-18 · Zhengyuan Yang, Jingen Liu, Jing Huang, Xiaodong He 외

In this study, we aim to predict the plausible future action steps given an observation of the past and study the task of instructional activity anticipation. Unlike previous anticipation tasks that aim at action label p…

Knowledge Distillation

Uncertainty-aware Action Decoupling Transformer for Action Anticipation

2024-01-01 · CVPR 2024 1 · Hongji Guo, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon Lee 외

Human action anticipation aims at predicting what people will do in the future based on past observations. In this paper we introduce Uncertainty-aware Action Decoupling Transformer (UADT) for action anticipation. Un…

Action Anticipation

Hierarchical GRU with Input-Conditioned Slot Queries for Ball Action Anticipation

2026-06-02 · Parthsarthi Rawat arxiv

We present a hierarchical model for ball action anticipation in football broadcast video. Given a 30-second observation window, the system predicts actions occurring in the subsequent 5-second window across 10 classes. A…

Action Anticipation

TransAction: ICL-SJTU Submission to EPIC-Kitchens Action Anticipation Challenge 2021

2021-07-28 · Xiao Gu, Jianing Qiu, Yao Guo, Benny Lo 외

In this report, the technical details of our submission to the EPIC-Kitchens Action Anticipation Challenge 2021 are given. We developed a hierarchical attention model for action anticipation, which leverages Transformer-…

Action Anticipation

What Would You Expect? Anticipating Egocentric Actions with Rolling-Unrolling LSTMs and Modality Attention

2019-05-22 · ICCV 2019 10 · Antonino Furnari, Giovanni Maria Farinella

Egocentric action anticipation consists in understanding which objects the camera wearer will interact with in the near future and which actions they will perform. We tackle the problem proposing an architecture able to …

Action AnticipationAction RecognitionEgocentric Activity RecognitionOptical Flow Estimation+2