paper-with-me

Papers

Advancing Video Self-Supervised Learning via Image Foundation Models

2025-05-25 · Jingwei Wu, Zhewei Huang, Chang Liu

In the past decade, image foundation models (IFMs) have achieved unprecedented progress. However, the potential of directly using IFMs for video self-supervised representation learning has largely been overlooked. In this study, we propose an advancing video self-supervised learning (AdViSe) approach, aimed at significantly reducing the training overhead of video representation models using pre-trained IFMs. Specifically, we first introduce temporal modeling modules (ResNet3D) to IFMs, constructing a video representation model. We then employ a video self-supervised learning approach, playback rate perception, to train temporal modules while freezing the IFM components. Experiments on UCF101 demonstrate that AdViSe achieves performance comparable to state-of-the-art methods while reducing training time by $3.4\times$ and GPU memory usage by $8.2\times$. This study offers fresh insights into low-cost video self-supervised learning based on pre-trained IFMs. Code is available at https://github.com/JingwWu/advise-video-ssl.

📄 PDF Abstract BibTeX arXiv:2505.19218

Code (1)

jingwwu/advise-video-ssl 공식 구현 pytorch

Tasks

GPURepresentation LearningSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Advancing Human Action Recognition with Foundation Models trained on Unlabeled Public Videos

2024-02-14 · Yang Qian, Yinan Sun, Ali Kargarandehkordi, Parnian Azizian 외

The increasing variety and quantity of tagged multimedia content on a variety of online platforms offer a unique opportunity to advance the field of human action recognition. In this study, we utilize 283,582 unique, unl…

Action RecognitionSelf-Supervised LearningTemporal Action Localization

A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources

2026-08-13 · Brunó B. Englert, Gijs Dubbelman arxiv

Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current scale required for pretraining visual foundation models may be unsustai…

Self-Supervised LearningRepresentation LearningImage ClassificationPose Estimation

VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models

2025-10-23 · Jesimon Barreto, Carlos Caetano, André Araujo, William Robson Schwartz arxiv

Foundation models have advanced computer vision by enabling strong performance across diverse tasks through large-scale pretraining and supervised fine-tuning. However, they may underperform in domains with distribution …

Self-Supervised Learning

Scaling up self-supervised learning for improved surgical foundation models

2025-01-16 · Tim J. M. Jaspers, Ronald L. P. D. de Jong, Yiping Li, Carolus H. J. Kusters 외

Foundation models have revolutionized computer vision by achieving vastly superior performance across diverse tasks through large-scale pretraining on extensive datasets. However, their application in surgical computer v…

Self-Supervised LearningSemantic Segmentation

LAVA: Language Audio Vision Alignment for Contrastive Video Pre-Training

2022-07-16 · Sumanth Gurram, Andy Fang, David Chan, John Canny

Generating representations of video data is of key importance in advancing the field of machine perception. Most current techniques rely on hand-annotated data, which can be difficult to work with, expensive to generate,…

Action RecognitionContrastive LearningTemporal Action Localization