Contrastive Learning of Global-Local Video Representations
Contrastive learning has delivered impressive results for various tasks in the self-supervised regime. However, existing approaches optimize for learning representations specific to downstream scenarios, i.e., \textit{global} representations suitable for tasks such as classification or \textit{local} representations for tasks such as detection and localization. While they produce satisfactory results in the intended downstream scenarios, they often fail to generalize to tasks that they were not originally designed for. In this work, we propose to learn video representations that generalize to both the tasks which require global semantic information (e.g., classification) and the tasks that require local fine-grained spatio-temporal information (e.g., localization). We achieve this by optimizing two contrastive objectives that together encourage our model to learn global-local visual information given audio signals. We show that the two objectives mutually improve the generalizability of the learned global-local representations, significantly outperforming their disjointly learned counterparts. We demonstrate our approach on various tasks including action/sound classification, lip reading, deepfake detection, event and sound localization (https://github.com/yunyikristy/global\_local).
Code (1)
Tasks
ClassificationContrastive LearningDeepFake DetectionFace SwappingGeneral ClassificationLip ReadingRepresentation LearningSound ClassificationSimilar Papers 제목 키워드 기반
Contrastive Learning of Global and Local Video Representations
Contrastive learning has delivered impressive results for various tasks in the self-supervised regime. However, existing approaches optimize for learning representations specific to downstream scenarios, i.e., global rep…
ClassificationContrastive LearningDeepFake DetectionFace Swapping+2Contrastive Self-Supervised Learning of Global-Local Audio-Visual Representations
Contrastive self-supervised learning has delivered impressive results in many audio-visual recognition tasks. However, existing approaches optimize for learning either global representations useful for high-level underst…
ClassificationDeepFake DetectionFace SwappingGeneral Classification+4TCLR: Temporal Contrastive Learning for Video Representation
Contrastive learning has nearly closed the gap between supervised and self-supervised learning of image representations, and has also been explored for videos. However, prior work on contrastive learning for video data h…
Action ClassificationAction RecognitionContrastive LearningGeneral Classification+6Self-Supervised Contrastive Learning for Videos using Differentiable Local Alignment
Robust frame-wise embeddings are essential to perform video analysis and understanding tasks. We present a self-supervised method for representation learning based on aligning temporal video sequences. Our framework uses…
Action RecognitionContrastive LearningRepresentation LearningVideo AlignmentMulti-Scale Contrastive Learning for Video Temporal Grounding
Temporal grounding, which localizes video moments related to a natural language query, is a core problem of vision-language learning and video understanding. To encode video moments of varying lengths, recent methods emp…
Contrastive LearningData AugmentationFormVideo Grounding+1