paper-with-me

홈 › Papers

Folded Recurrent Neural Networks for Future Video Prediction

2017-12-01 · ECCV 2018 9 · Marc Oliu, Javier Selva, Sergio Escalera

Future video prediction is an ill-posed Computer Vision problem that recently received much attention. Its main challenges are the high variability in video content, the propagation of errors through time, and the non-specificity of the future frames: given a sequence of past frames there is a continuous distribution of possible futures. This work introduces bijective Gated Recurrent Units, a double mapping between the input and output of a GRU layer. This allows for recurrent auto-encoders with state sharing between encoder and decoder, stratifying the sequence representation and helping to prevent capacity problems. We show how with this topology only the encoder or decoder needs to be applied for input encoding and prediction, respectively. This reduces the computational cost and avoids re-encoding the predictions when generating a sequence of frames, mitigating the propagation of errors. Furthermore, it is possible to remove layers from an already trained model, giving an insight to the role performed by each layer and making the model more explainable. We evaluate our approach on three video datasets, outperforming state of the art prediction results on MMNIST and UCF101, and obtaining competitive results on KTH with 2 and 3 times less memory usage and computational cost than the best scored approach.

📄 PDF Abstract BibTeX arXiv:1712.00311

Code (1)

moliusimon/frnn 공식 구현 tf

Tasks

DecoderPredictionSpecificityVideo Prediction

Methods 이 논문이 사용한 방법론

GRU A Gated Recurrent Unit, or GRU, is a type of recurrent neural network. It is similar to an LSTM, but only has two gates - a reset…

Similar Papers 제목 키워드 기반

Long History Short-Term Memory for Long-Term Video Prediction

2019-09-25 · Wonmin Byeon, Jan Kautz

While video prediction approaches have advanced considerably in recent years, learning to predict long-term future is challenging — ambiguous future or error propagation over time yield blurry predictions. To address thi…

Video Prediction

Predcnn: Predictive learning with cascade convolutions

2018-07-01 · Twenty-Seventh International Joint Conference on Artificial Intelligence {IJCAI-18} 2018 7 · Ziru Xu, Yunbo Wang, Mingsheng Long, Jian-Min Wang

Predicting future frames in videos remains an unsolved but challenging problem. Mainstream recurrent models suffer from huge memory usage and computation cost, while convolutional models are unable to effectively capture…

Pose PredictionVideo Prediction

Local Frequency Domain Transformer Networks for Video Prediction

2021-05-10 · Hafez Farazi, Jan Nogga, Sven Behnke

Video prediction is commonly referred to as forecasting future frames of a video sequence provided several past frames thereof. It remains a challenging domain as visual scenes evolve according to complex underlying dyna…

Motion SegmentationPredictionVideo Prediction

A Log-likelihood Regularized KL Divergence for Video Prediction with A 3D Convolutional Variational Recurrent Network

2020-12-11 · Haziq Razali, Basura Fernando

The use of latent variable models has shown to be a powerful tool for modeling probability distributions over sequences. In this paper, we introduce a new variational model that extends the recurrent network in two ways …

PredictionVideo Prediction

Future semantic segmentation of time-lapsed videos with large temporal displacement

2018-12-27 · Talha Siddiqui, Samarth Bharadwaj

An important aspect of video understanding is the ability to predict the evolution of its content in the future. This paper presents a future frame semantic segmentation technique for predicting semantic masks of the cur…

SegmentationSemantic SegmentationVideo Understanding