paper-with-me

Papers

Learning Temporal Sentence Grounding From Narrated EgoVideos

2023-10-26 · Kevin Flanagan, Dima Damen, Michael Wray

The onset of long-form egocentric datasets such as Ego4D and EPIC-Kitchens presents a new challenge for the task of Temporal Sentence Grounding (TSG). Compared to traditional benchmarks on which this task is evaluated, these datasets offer finer-grained sentences to ground in notably longer videos. In this paper, we develop an approach for learning to ground sentences in these datasets using only narrations and their corresponding rough narration timestamps. We propose to artificially merge clips to train for temporal grounding in a contrastive manner using text-conditioning attention. This Clip Merging (CliMer) approach is shown to be effective when compared with a high performing TSG method -- e.g. mean R@1 improves from 3.9 to 5.7 on Ego4D and from 10.7 to 13.0 on EPIC-Kitchens. Code and data splits available from: https://github.com/keflanagan/CliMer

📄 PDF Abstract BibTeX arXiv:2310.17395

Code (1)

keflanagan/climer 공식 구현 pytorch

Tasks

SentenceTemporal Sentence Grounding

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Self-view Grounding Given a Narrated 360° Video

2017-11-23 · Shih-Han Chou, Yi-Chun Chen, Kuo-Hao Zeng, Hou-Ning Hu 외

Narrated 360{\deg} videos are typically provided in many touring scenarios to mimic real-world experience. However, previous work has shown that smart assistance (i.e., providing visual guidance) can significantly help u…

SentenceVisual Grounding

What When and Where? Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions

2024-01-01 · CVPR 2024 1 · Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann 외

Spatio-temporal grounding describes the task of localizing events in space and time e.g. in video data based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bo…

Representation Learning

What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions

2023-03-29 · Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann 외

Spatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bou…

Representation LearningSpatio-Temporal Video Grounding

Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning

2023-02-27 · CVPR 2023 1 · Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech 외

In this work, we introduce Vid2Seq, a multi-modal single-stage dense event captioning model pretrained on narrated videos which are readily-available at scale. The Vid2Seq architecture augments a language model with spec…

Dense Video CaptioningLanguage ModelingLanguage ModellingSentence+1

When Did It Happen? Duration-informed Temporal Localization of Narrated Actions in Vlogs

2022-02-16 · Oana Ignat, Santiago Castro, YuHang Zhou, Jiajun Bao 외

We consider the task of temporal human action localization in lifestyle vlogs. We introduce a novel dataset consisting of manual annotations of temporal localization for 13,000 narrated actions in 1,200 video clips. We p…

Action LocalizationTemporal Action LocalizationTemporal Localization