paper-with-me

홈 › Papers

Fast Fourier Inception Networks for Occluded Video Prediction

2023-06-17 · Ping Li, Chenhan Zhang, Xianghua Xu

Video prediction is a pixel-level task that generates future frames by employing the historical frames. There often exist continuous complex motions, such as object overlapping and scene occlusion in video, which poses great challenges to this task. Previous works either fail to well capture the long-term temporal dynamics or do not handle the occlusion masks. To address these issues, we develop the fully convolutional Fast Fourier Inception Networks for video prediction, termed \textit{FFINet}, which includes two primary components, \ie, the occlusion inpainter and the spatiotemporal translator. The former adopts the fast Fourier convolutions to enlarge the receptive field, such that the missing areas (occlusion) with complex geometric structures are filled by the inpainter. The latter employs the stacked Fourier transform inception module to learn the temporal evolution by group convolutions and the spatial movement by channel-wise Fourier convolutions, which captures both the local and the global spatiotemporal features. This encourages generating more realistic and high-quality future frames. To optimize the model, the recovery loss is imposed to the objective, \ie, minimizing the mean square error between the ground-truth frame and the recovery frame. Both quantitative and qualitative experimental results on five benchmarks, including Moving MNIST, TaxiBJ, Human3.6M, Caltech Pedestrian, and KTH, have demonstrated the superiority of the proposed approach. Our code is available at GitHub.

📄 PDF Abstract BibTeX arXiv:2306.10346

Code (1)

mlvccn/research 공식 구현 pytorch

Tasks

PredictionVideo Prediction

Methods 이 논문이 사용한 방법론

fail 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Inception Module An Inception Module is an image model block that aims to approximate an optimal local sparse structure in a CNN. Put simply, it allows for us to use multiple types of filter…

Similar Papers 제목 키워드 기반

Inception-inspired LSTM for Next-frame Video Prediction

2019-08-28 · Matin Hosseini, Anthony S. Maida, Majid Hosseini, Gottumukkala Raju

The problem of video frame prediction has received much interest due to its relevance to many computer vision applications such as autonomous vehicles or robotics. Supervised methods for video frame prediction rely on la…

Autonomous Vehiclesimage-classificationImage ClassificationPrediction+1

A Comparison of Deep Learning Models for the Prediction of Hand Hygiene Videos

2021-11-03 · Rashmi Bakshi

This paper presents a comparison of various deep learning models such as Exception, Resnet-50, and Inception V3 for the classification and prediction of hand hygiene gestures, which were recorded in accordance with the W…

GPU

Disentangling Propagation and Generation for Video Prediction

2018-12-02 · ICCV 2019 10 · Hang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang 외

A dynamic scene has two types of elements: those that move fluidly and can be predicted from previous frames, and those which are disoccluded (exposed) and cannot be extrapolated. Prior approaches to video prediction typ…

Predict Future Video FramesPredictionVideo Prediction

Adversarial Video Generation on Complex Datasets

2019-07-15 · Aidan Clark, Jeff Donahue, Karen Simonyan

Generative models of natural images have progressed towards high fidelity samples by the strong leveraging of scale. We attempt to carry this success to the field of video modeling by showing that large Generative Advers…

3D Character Animation From A Single PhotoVideo GenerationVideo Prediction

Flow and Depth Assisted Video Prediction with Latent Transformer

2025-11-20 · Eliyas Suleyman, Paul Henderson, Eksan Firkat, Nicolas Pugeault arxiv

Video prediction is a fundamental task for various downstream applications, including robotics and world modeling. Although general video prediction models have achieved remarkable performance in standard scenarios, occl…

Video Prediction