paper-with-me

홈 › Papers

Modeling Video As Stochastic Processes for Fine-Grained Video Representation Learning

2023-01-01 · CVPR 2023 1 · Heng Zhang, Daqing Liu, Qi Zheng, Bing Su

A meaningful video is semantically coherent and changes smoothly. However, most existing fine-grained video representation learning methods learn frame-wise features by aligning frames across videos or exploring relevance between multiple views, neglecting the inherent dynamic process of each video. In this paper, we propose to learn video representations by modeling Video as Stochastic Processes (VSP) via a novel process-based contrastive learning framework, which aims to discriminate between video processes and simultaneously capture the temporal dynamics in the processes. Specifically, we enforce the embeddings of the frame sequence of interest to approximate a goal-oriented stochastic process, i.e., Brownian bridge, in the latent space via a process-based contrastive loss. To construct the Brownian bridge, we adapt specialized sampling strategies under different annotations for both self-supervised and weakly-supervised learning. Experimental results on four datasets show that VSP stands as a state-of-the-art method for various video understanding tasks, including phase progression, phase classification and frame retrieval. Code is available at 'https://github.com/hengRUC/VSP'.

📄 PDF Abstract BibTeX

Code (1)

hengruc/vsp 공식 구현 pytorch

Tasks

Contrastive LearningRepresentation LearningRetrievalVideo UnderstandingWeakly-supervised Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

2026-04-09 · Handong Li, Zikang Liu, Longteng Guo, Tongtian Yue 외 arxiv

Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception through irreversible information disposal or i…

Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models

2024-12-24 · CVPR 2025 1 · Jinhui Yi, Syed Talal Wasim, Yanan Luo, Muzammal Naseer 외

We present an efficient encoder-free approach for video-language understanding that achieves competitive performance while significantly reducing computational overhead. Current video-language models typically rely on he…

Question AnsweringVideo Question Answering

Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling

2022-10-08 · Hsin-Ying Lee, Hung-Ting Su, Bing-Chen Tsai, Tsung-Han Wu 외

While recent large-scale video-language pre-training made great progress in video question answering, the design of spatial modeling of video-language models is less fine-grained than that of image-language models; exist…

Language ModelingLanguage ModellingQuestion AnsweringVideo Question Answering

Fine-grained Audible Video Description

2023-03-27 · CVPR 2023 1 · Xuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin 외

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and s…

Language ModelingLanguage ModellingMasked Language ModelingSentence+4

Fine-Grained Video Captioning for Sports Narrative

2018-06-01 · CVPR 2018 6 · Huanyu Yu, Shuo Cheng, Bingbing Ni, Minsi Wang 외

Despite recent emergence of video caption methods, how to generate fine-grained video descriptions (i.e., long and detailed commentary about individual movements of multiple subjects as well as their frequent interaction…

2kVideo Captioning