MoQuad: Motion-focused Quadruple Construction for Video Contrastive Learning
Learning effective motion features is an essential pursuit of video representation learning. This paper presents a simple yet effective sample construction strategy to boost the learning of motion features in video contrastive learning. The proposed method, dubbed Motion-focused Quadruple Construction (MoQuad), augments the instance discrimination by meticulously disturbing the appearance and motion of both the positive and negative samples to create a quadruple for each video instance, such that the model is encouraged to exploit motion information. Unlike recent approaches that create extra auxiliary tasks for learning motion features or apply explicit temporal modelling, our method keeps the simple and clean contrastive learning paradigm (i.e.,SimCLR) without multi-task learning or extra modelling. In addition, we design two extra training strategies by analyzing initial MoQuad experiments. By simply applying MoQuad to SimCLR, extensive experiments show that we achieve superior performance on downstream tasks compared to the state of the arts. Notably, on the UCF-101 action recognition task, we achieve 93.7% accuracy after pre-training the model on Kinetics-400 for only 200 epochs, surpassing various previous methods
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionContrastive LearningMulti-Task LearningRepresentation LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
ECQED: Emotion-Cause Quadruple Extraction in Dialogs
The existing emotion-cause pair extraction (ECPE) task, unfortunately, ignores extracting the emotion type and cause type, while these fine-grained meta-information can be practically useful in real-world applications, i…
Emotion-Cause Pair ExtractionVideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors
Generative methods for image and video editing use generative models as priors to perform edits despite incomplete information, such as changing the composition of 3D objects shown in a single image. Recent methods have …
3D ReconstructionObjectVideo EditingLocoMotion: Learning Motion-Focused Video-Language Representations
This paper strives for motion-focused video-language representations. Existing methods to learn video-language representations use spatial-focused data, where identifying the objects and scene is often enough to distingu…
3D Reconstruction of Whole Stomach from Endoscope Video Using Structure-from-Motion
Gastric endoscopy is a common clinical practice that enables medical doctors to diagnose the stomach inside a body. In order to identify a gastric lesion's location such as early gastric cancer within the stomach, this w…
3D Reconstructionchannel selectionVideo Anomaly Detection By The Duality Of Normality-Granted Optical Flow
Video anomaly detection is a challenging task because of diverse abnormal events. To this task, methods based on reconstruction and prediction are wildly used in recent works, which are built on the assumption that learn…
Anomaly DetectionOptical Flow EstimationPredictionVideo Anomaly Detection