Combining Contrastive and Supervised Learning for Video Super-Resolution Detection
Upscaled video detection is a helpful tool in multimedia forensics, but it is a challenging task that involves various upscaling and compression algorithms. There are many resolution-enhancement methods, including interpolation and deep-learning-based super-resolution, and they leave unique traces. In this work, we propose a new upscaled-resolution-detection method based on learning of visual representations using contrastive and cross-entropy losses. To explain how the method detects videos, we systematically review the major components of our framework - in particular, we show that most data-augmentation approaches hinder the learning of the method. Through extensive experiments on various datasets, we demonstrate that our method effectively detects upscaling even in compressed videos and outperforms the state-of-the-art alternatives. The code and models are publicly available at https://github.com/msu-video-group/SRDM
Code (1)
Tasks
Data AugmentationSuper-ResolutionVideo Super-ResolutionSimilar Papers 제목 키워드 기반
Pixel-level Counterfactual Contrastive Learning for Medical Image Segmentation
Image segmentation relies on large annotated datasets, which are expensive and slow to produce. Silver-standard (AI-generated) labels are easier to obtain, but they risk introducing bias. Self-supervised learning, needin…
Medical Image SegmentationSelf-Supervised LearningRepresentation LearningContrastive LearningSelf-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution
Existing diffusion-based video super-resolution (VSR) methods are susceptible to introducing complex degradations and noticeable artifacts into high-resolution videos due to their inherent randomness. In this paper, …
Contrastive LearningMambaSelf-Supervised LearningSuper-Resolution+1Self-supervised and Weakly Supervised Contrastive Learning for Frame-wise Action Representations
Previous work on action representation learning focused on global representations for short video clips. In contrast, many practical applications, such as video alignment, strongly demand learning the intensive represent…
Action ClassificationContrastive LearningRepresentation LearningRetrieval+23D Human Pose, Shape and Texture from Low-Resolution Images and Videos
3D human pose and shape estimation from monocular images has been an active research area in computer vision. Existing deep learning methods for this task rely on high-resolution input, which however, is not always avail…
3D human pose and shape estimationContrastive LearningSuper-ResolutionPretext-Contrastive Learning: Toward Good Practices in Self-supervised Video Representation Leaning
Recently, pretext-task based methods are proposed one after another in self-supervised video feature learning. Meanwhile, contrastive learning methods also yield good performance. Usually, new methods can beat previous o…
Contrastive LearningData AugmentationSelf-Supervised Action RecognitionSelf-Supervised Learning+2