paper-with-me

홈 › Papers

Learning by Aligning Videos in Time

2021-03-31 · CVPR 2021 1 · Sanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed, Andrey Konin, Muhammad Zeeshan Zia, Quoc-Huy Tran

We present a self-supervised approach for learning video representations using temporal video alignment as a pretext task, while exploiting both frame-level and video-level information. We leverage a novel combination of temporal alignment loss and temporal regularization terms, which can be used as supervision signals for training an encoder network. Specifically, the temporal alignment loss (i.e., Soft-DTW) aims for the minimum cost for temporally aligning videos in the embedding space. However, optimizing solely for this term leads to trivial solutions, particularly, one where all frames get mapped to a small cluster in the embedding space. To overcome this problem, we propose a temporal regularization term (i.e., Contrastive-IDM) which encourages different frames to be mapped to different points in the embedding space. Extensive evaluations on various tasks, including action phase classification, action phase progression, and fine-grained frame retrieval, on three datasets, namely Pouring, Penn Action, and IKEA ASM, show superior performance of our approach over state-of-the-art methods for self-supervised representation learning from videos. In addition, our method provides significant performance gain where labeled data is lacking. Our code and labels are available on our research website: https://retrocausal.ai/research/

📄 PDF Abstract BibTeX arXiv:2103.17260

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningRetrievalVideo Alignment

Similar Papers 제목 키워드 기반

LAMV: Learning to Align and Match Videos With Kernelized Temporal Layers

2018-06-01 · CVPR 2018 6 · Lorenzo Baraldi, Matthijs Douze, Rita Cucchiara, Hervé Jégou

This paper considers a learnable approach for comparing and aligning videos. Our architecture builds upon and revisits temporal match kernels within neural networks: we propose a new temporal layer that finds temporal al…

Copy DetectionRetrievalTripletVideo Alignment+1

Aligning Videos in Space and Time

2020-07-09 · ECCV 2020 8 · Senthil Purushwalkam, Tian Ye, Saurabh Gupta, Abhinav Gupta

In this paper, we focus on the task of extracting visual correspondences across videos. Given a query video clip from an action class, we aim to align it with training videos in space and time. Obtaining training data fo…

HT-Step: Aligning Instructional Articles with How-To Videos

2023-09-26 · NeurIPS 2023 11

We introduce HT-Step, a large-scale dataset containing temporal annotations of instructional article steps in cooking videos. It includes 122k segment-level annotations over 20k narrated videos (approximately 2.3k hours)…

Weakly-Supervised Alignment of Video With Text

2015-05-22 · ICCV 2015 12 · Piotr Bojanowski, Rémi Lajugie, Edouard Grave, Francis Bach 외

Suppose that we are given a set of videos, along with natural language descriptions in the form of multiple sentences (e.g., manual annotations, movie scripts, sport summaries etc.), and that these sentences appear in th…

Sentence

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

2025-09-10 · Garry Yang, Zizhe Chen, Man Hon Wong, Haoyu Lei 외 arxiv

Large Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic video content. Despite their progress, LVMs…