paper-with-me

Papers

MoQuad: Motion-focused Quadruple Construction for Video Contrastive Learning

2022-12-21 · YuAn Liu, Jiacheng Chen, Hao Wu

Learning effective motion features is an essential pursuit of video representation learning. This paper presents a simple yet effective sample construction strategy to boost the learning of motion features in video contrastive learning. The proposed method, dubbed Motion-focused Quadruple Construction (MoQuad), augments the instance discrimination by meticulously disturbing the appearance and motion of both the positive and negative samples to create a quadruple for each video instance, such that the model is encouraged to exploit motion information. Unlike recent approaches that create extra auxiliary tasks for learning motion features or apply explicit temporal modelling, our method keeps the simple and clean contrastive learning paradigm (i.e.,SimCLR) without multi-task learning or extra modelling. In addition, we design two extra training strategies by analyzing initial MoQuad experiments. By simply applying MoQuad to SimCLR, extensive experiments show that we achieve superior performance on downstream tasks compared to the state of the arts. Notably, on the UCF-101 action recognition task, we achieve 93.7% accuracy after pre-training the model on Kinetics-400 for only 200 epochs, surpassing various previous methods

📄 PDF Abstract BibTeX arXiv:2212.10870

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionContrastive LearningMulti-Task LearningRepresentation Learning

Methods 이 논문이 사용한 방법론

Bitcoin Customer Service Number +1-833-534-1729 설명 없음
Average Pooling 설명 없음
Batch Normalization 설명 없음
Residual Block Residual Blocks are skip-connection blocks that learn residual functions with reference to the layer inputs, instead of learning unreferenced functions. They were introduced…
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Kaiming Initialization 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

ECQED: Emotion-Cause Quadruple Extraction in Dialogs

2023-06-06 · Li Zheng, Donghong Ji, Fei Li, Hao Fei 외

The existing emotion-cause pair extraction (ECPE) task, unfortunately, ignores extracting the emotion type and cause type, while these fine-grained meta-information can be practically useful in real-world applications, i…

Emotion-Cause Pair Extraction

VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors

2025-03-03 · CVPR 2025 1 · Juil Koo, Paul Guerrero, Chun-Hao Paul Huang, Duygu Ceylan 외

Generative methods for image and video editing use generative models as priors to perform edits despite incomplete information, such as changing the composition of 3D objects shown in a single image. Recent methods have …

3D ReconstructionObjectVideo Editing

LocoMotion: Learning Motion-Focused Video-Language Representations

2024-10-15 · Hazel Doughty, Fida Mohammad Thoker, Cees G. M. Snoek

This paper strives for motion-focused video-language representations. Existing methods to learn video-language representations use spatial-focused data, where identifying the objects and scene is often enough to distingu…

3D Reconstruction of Whole Stomach from Endoscope Video Using Structure-from-Motion

2019-05-30 · Aji Resindra Widya, Yusuke Monno, Kosuke Imahori, Masatoshi Okutomi 외

Gastric endoscopy is a common clinical practice that enables medical doctors to diagnose the stomach inside a body. In order to identify a gastric lesion's location such as early gastric cancer within the stomach, this w…

3D Reconstructionchannel selection

Video Anomaly Detection By The Duality Of Normality-Granted Optical Flow

2021-05-10 · Hongyong Wang, Xinjian Zhang, Su Yang, Weishan Zhang

Video anomaly detection is a challenging task because of diverse abnormal events. To this task, methods based on reconstruction and prediction are wildly used in recent works, which are built on the assumption that learn…

Anomaly DetectionOptical Flow EstimationPredictionVideo Anomaly Detection