Video-adverb retrieval with compositional adverb-action embeddings
Retrieving adverbs that describe an action in a video poses a crucial step towards fine-grained video understanding. We propose a framework for video-to-adverb retrieval (and vice versa) that aligns video embeddings with their matching compositional adverb-action text embedding in a joint embedding space. The compositional adverb-action text embedding is learned using a residual gating mechanism, along with a novel training objective consisting of triplet losses and a regression target. Our method achieves state-of-the-art performance on five recent benchmarks for video-adverb retrieval. Furthermore, we introduce dataset splits to benchmark video-adverb retrieval for unseen adverb-action compositions on subsets of the MSR-VTT Adverbs and ActivityNet Adverbs datasets. Our proposed framework outperforms all prior works for the generalisation task of retrieving adverbs from videos for unseen adverb-action compositions. Code and dataset splits are available at https://hummelth.github.io/ReGaDa/.
Code (1)
Tasks
TripletVideo-Adverb RetrievalVideo-Adverb Retrieval (Unseen Compositions)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Action Modifiers: Learning from Adverbs in Instructional Videos
We present a method to learn a representation for adverbs from instructional videos using weak supervision from the accompanying narrations. Key to our method is the fact that the visual representation of the adverb is h…
Video-Adverb RetrievalHow Do You Do It? Fine-Grained Action Understanding with Pseudo-Adverbs
We aim to understand how actions are performed and identify subtle differences, such as 'fold firmly' vs. 'fold gently'. To this end, we propose a method which recognizes adverbs across different actions. However, such f…
Video-Adverb RetrievalVideo-Adverb Retrieval (Unseen Compositions)Learning Action Changes by Measuring Verb-Adverb Textual Relationships
The goal of this work is to understand the way actions are performed in videos. That is, given a video, we aim to predict an adverb indicating a modification applied to the action (e.g. cut "finely"). We cast this proble…
Video-Adverb RetrievalHuman Action Adverb Recognition: ADHA Dataset and A Three-Stream Hybrid Model
We introduce the first benchmark for a new problem --- recognizing human action adverbs (HAA): "Adverbs Describing Human Actions" (ADHA). This is the first step for computer vision to change over from pattern recognition…
Action RecognitionImage CaptioningTemporal Action LocalizationReasoning over the Behaviour of Objects in Video-Clips for Adverb-Type Recognition
In this work, following the intuition that adverbs describing scene-sequences are best identified by reasoning over high-level concepts of object-behavior, we propose the design of a new framework that reasons over objec…
Object