When Did It Happen? Duration-informed Temporal Localization of Narrated Actions in Vlogs
We consider the task of temporal human action localization in lifestyle vlogs. We introduce a novel dataset consisting of manual annotations of temporal localization for 13,000 narrated actions in 1,200 video clips. We present an extensive analysis of this data, which allows us to better understand how the language and visual modalities interact throughout the videos. We propose a simple yet effective method to localize the narrated actions based on their expected duration. Through several experiments and analyses, we show that our method brings complementary information with respect to previous methods, and leads to improvements over previous work for the task of temporal action localization.
Code (1)
Tasks
Action LocalizationTemporal Action LocalizationTemporal LocalizationSimilar Papers 제목 키워드 기반
Revisiting Anchor Mechanisms for Temporal Action Localization
Most of the current action localization methods follow an anchor-based pipeline: depicting action instances by pre-defined anchors, learning to select the anchors closest to the ground truth, and predicting the confidenc…
Action LocalizationTemporal Action LocalizationDual Attention Matching for Audio-Visual Event Localization
In this paper, we investigate the audio-visual event localization problem. This task is to localize a visible and audible event in a video. Previous methods first divide a video into short segments, and then fuse visual …
audio-visual event localizationEncode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization
Audio language models process input audio into rich frame-level representations, but the standard approach to temporal localization generates timestamps as sequences of text tokens, which discards the frame-level represe…
Speaker DiarizationWord AlignmentScale Matters: Temporal Scale Aggregation Network for Precise Action Localization in Untrimmed Videos
Temporal action localization is a recently-emerging task, aiming to localize video segments from untrimmed videos that contain specific actions. Despite the remarkable recent progress, most two-stage action localization …
Action LocalizationTemporal Action LocalizationRethinking the Faster R-CNN Architecture for Temporal Action Localization
We propose TAL-Net, an improved approach to temporal action localization in video that is inspired by the Faster R-CNN object detection framework. TAL-Net addresses three key shortcomings of existing approaches: (1) we i…
Action ClassificationAction LocalizationGeneral Classificationobject-detection+2