A Hierarchical Variational Neural Uncertainty Model for Stochastic Video Prediction
Predicting the future frames of a video is a challenging task, in part due to the underlying stochastic real-world phenomena. Prior approaches to solve this task typically estimate a latent prior characterizing this stochasticity, however do not account for the predictive uncertainty of the (deep learning) model. Such approaches often derive the training signal from the mean-squared error (MSE) between the generated frame and the ground truth, which can lead to sub-optimal training, especially when the predictive uncertainty is high. Towards this end, we introduce Neural Uncertainty Quantifier (NUQ) - a stochastic quantification of the model's predictive uncertainty, and use it to weigh the MSE loss. We propose a hierarchical, variational framework to derive NUQ in a principled manner using a deep, Bayesian graphical model. Our experiments on four benchmark stochastic video prediction datasets show that our proposed framework trains more effectively compared to the state-of-the-art models (especially when the training sets are small), while demonstrating better video generation quality and diversity against several evaluation metrics.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityVideo GenerationVideo PredictionSimilar Papers 제목 키워드 기반
Doubly Stochastic Variational Inference for Neural Processes with Hierarchical Latent Variables
Neural processes (NPs) constitute a family of variational approximate models for stochastic processes with promising properties in computational efficiency and uncertainty quantification. These processes use neural netwo…
Computational EfficiencyregressionUncertainty QuantificationVariational InferenceDeep Variational Luenberger-type Observer for Stochastic Video Prediction
Considering the inherent stochasticity and uncertainty, predicting future video frames is exceptionally challenging. In this work, we study the problem of video prediction by combining interpretability of stochastic stat…
Representation LearningState Space ModelsVideo PredictionVocal Bursts Type PredictionGreedy Hierarchical Variational Autoencoders for Large-Scale Video Prediction
A video prediction model that generalizes to diverse scenes would enable intelligent agents such as robots to perform a variety of tasks via planning with the model. However, while existing video prediction models have p…
PredictionVideo PredictionLearning to Generate Videos Using Neural Uncertainty Priors
Predicting the future frames of a video is a challenging task, in part due to the underlying stochastic real-world phenomena. Previous approaches attempt to tackle this issue by estimating a latent prior characterizing t…
DiversityVideo GenerationStochastic Variational Video Prediction
Predicting the future in real-world settings, particularly from raw sensory observations such as images, is exceptionally challenging. Real-world events can be stochastic and unpredictable, and the high dimensionality an…
PredictionVideo GenerationVideo Prediction