State-space Decomposition Model for Video Prediction Considering Long-term Motion Trend
Stochastic video prediction enables the consideration of uncertainty in future motion, thereby providing a better reflection of the dynamic nature of the environment. Stochastic video prediction methods based on image auto-regressive recurrent models need to feed their predictions back into the latent space. Conversely, the state-space models, which decouple frame synthesis and temporal prediction, proves to be more efficient. However, inferring long-term temporal information about motion and generalizing to dynamic scenarios under non-stationary assumptions remains an unresolved challenge. In this paper, we propose a state-space decomposition stochastic video prediction model that decomposes the overall video frame generation into deterministic appearance prediction and stochastic motion prediction. Through adaptive decomposition, the model's generalization capability to dynamic scenarios is enhanced. In the context of motion prediction, obtaining a prior on the long-term trend of future motion is crucial. Thus, in the stochastic motion prediction branch, we infer the long-term motion trend from conditional frames to guide the generation of future frames that exhibit high consistency with the conditional frames. Experimental results demonstrate that our model outperforms baselines on multiple datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
motion predictionPredictionState Space ModelsVideo PredictionSimilar Papers 제목 키워드 기반
Deep Variational Luenberger-type Observer for Stochastic Video Prediction
Considering the inherent stochasticity and uncertainty, predicting future video frames is exceptionally challenging. In this work, we study the problem of video prediction by combining interpretability of stochastic stat…
Representation LearningState Space ModelsVideo PredictionVocal Bursts Type PredictionUnsupervised object-centric video generation and decomposition in 3D
A natural approach to generative modeling of videos is to represent them as a composition of moving objects. Recent works model a set of 2D sprites over a slowly-varying background, but without considering the underlying…
3D Object DetectionDepth EstimationDepth PredictionInstance Segmentation+4Unifying Theorems for Subspace Identification and Dynamic Mode Decomposition
This paper presents unifying results for subspace identification (SID) and dynamic mode decomposition (DMD) for autonomous dynamical systems. We observe that SID seeks to solve an optimization problem to estimate an exte…
On the Benefits of Instance Decomposition in Video Prediction Models
Video prediction is a crucial task for intelligent agents such as robots and autonomous vehicles, since it enables them to anticipate and act early on time-critical incidents. State-of-the-art video prediction methods ty…
Autonomous VehiclesPredictionVideo PredictionTruck Parking Usage Prediction with Decomposed Graph Neural Networks
Truck parking on freight corridors faces the major challenge of insufficient parking spaces. This is exacerbated by the Hour-of-Service (HOS) regulations, which often result in unauthorized parking practices, causing saf…
Graph Neural NetworkPrediction