paper-with-me

홈 › Papers

Flexible Spatio-Temporal Networks for Video Prediction

2017-07-01 · CVPR 2017 7 · Chaochao Lu, Michael Hirsch, Bernhard Scholkopf

We describe a modular framework for video frame prediction. We refer to it as a Flexible Spatio-Temporal Network (FSTN) as it allows the extrapolation of a video sequence as well as the estimation of synthetic frames lying in between observed frames and thus the generation of slow-motion videos. By devising a customized objective function comprising decoding, encoding, and adversarial losses, we are able to mitigate the common problem of blurry predictions, managing to retain high frequency information even for relatively distant future predictions. We propose and analyse different training strategies to optimize our model. Extensive experiments on several challenging public datasets demonstrate both the versatility and validity of our model.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

PredictionVideo Prediction

Similar Papers 제목 키워드 기반

MotionRNN: A Flexible Model for Video Prediction with Spacetime-Varying Motions

2021-03-03 · CVPR 2021 1 · Haixu Wu, Zhiyu Yao, Jianmin Wang, Mingsheng Long

This paper tackles video prediction from a new dimension of predicting spacetime-varying motions that are incessantly changing across both space and time. Prior methods mainly capture the temporal state transitions but o…

PredictionVideo Prediction

STAU: A SpatioTemporal-Aware Unit for Video Prediction and Beyond

2022-04-20 · Zheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma 외

Video prediction aims to predict future frames by modeling the complex spatiotemporal dynamics in videos. However, most of the existing methods only model the temporal information and the spatial information for videos i…

Action Recognitionobject-detectionObject DetectionPrediction+1

StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales

2025-10-17 · Nyle Siddiqui, Rohit Gupta, Sirnam Swetha, Mubarak Shah arxiv

State space models (SSMs) have emerged as a competitive alternative to transformers in various tasks. Their linear complexity and hidden-state recurrence make them particularly attractive for modeling long sequences, whe…

Action Recognition

TiS-TSL: Image-Label Supervised Surgical Video Stereo Matching via Time-Switchable Teacher-Student Learning

2025-11-10 · Rui Wang, Ying Zhou, Hao Wang, Wenwei Zhang 외 arxiv

Stereo matching in minimally invasive surgery (MIS) is essential for next-generation navigation and augmented reality. Yet, dense disparity supervision is nearly impossible due to anatomical constraints, typically limiti…

Deepfake Video Detection with Spatiotemporal Dropout Transformer

2022-07-14 · Daichi Zhang, Fanzhao Lin, Yingying Hua, Pengju Wang 외

While the abuse of deepfake technology has caused serious concerns recently, how to detect deepfake videos is still a challenge due to the high photo-realistic synthesis of each frame. Existing image-level approaches oft…

Data AugmentationFace Swapping