paper-with-me

Papers

Stochastic Dynamics for Video Infilling

2018-09-01 · Qiangeng Xu, Hanwang Zhang, Weiyue Wang, Peter N. Belhumeur, Ulrich Neumann

In this paper, we introduce a stochastic dynamics video infilling (SDVI) framework to generate frames between long intervals in a video. Our task differs from video interpolation which aims to produce transitional frames for a short interval between every two frames and increase the temporal resolution. Our task, namely video infilling, however, aims to infill long intervals with plausible frame sequences. Our framework models the infilling as a constrained stochastic generation process and sequentially samples dynamics from the inferred distribution. SDVI consists of two parts: (1) a bi-directional constraint propagation module to guarantee the spatial-temporal coherence among frames, (2) a stochastic sampling process to generate dynamics from the inferred distributions. Experimental results show that SDVI can generate clear frame sequences with varying contents. Moreover, motions in the generated sequence are realistic and able to transfer smoothly from the given start frame to the terminal frame. Our project site is https://xharlie.github.io/projects/project_sites/SDVI/video_results.html

📄 PDF Abstract BibTeX arXiv:1809.00263

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SAGA: Stochastic Whole-Body Grasping with Contact

2021-12-19 · Yan Wu, Jiahao Wang, Yan Zhang, Siwei Zhang 외

The synthesis of human grasping has numerous applications including AR/VR, video games and robotics. While methods have been proposed to generate realistic hand-object interaction for object grasping and manipulation, th…

Object

Revealing Disocclusions in Temporal View Synthesis through Infilling Vector Prediction

2021-10-17 · Vijayalakshmi Kanchana, Nagabhushan Somraj, Suraj Yadwad, Rajiv Soundararajan

We consider the problem of temporal view synthesis, where the goal is to predict a future video frame from the past frames using knowledge of the depth and relative camera motion. In contrast to revealing the disoccluded…

Temporal View Synthesis

MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling

2024-05-28 · BoWen Zhang, Xiaofei Xie, Haotian Lu, Na Ma 외

Diffusion-based video generation has achieved significant progress, yet generating multiple actions that occur sequentially remains a formidable task. Directly generating a video with sequential actions can be extremely …

Video Generation

Diffusion Models for Video Prediction and Infilling

2022-06-15 · Tobias Höppe, Arash Mehrjou, Stefan Bauer, Didrik Nielsen 외

Predicting and anticipating future outcomes or reasoning about missing information in a sequence are critical skills for agents to be able to make intelligent decisions. This requires strong, temporally coherent generati…

PredictionVideo GenerationVideo Prediction

Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

2023-05-23 · Vaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He 외

Despite exciting recent results showing vision-language systems' capacity to reason about images using natural language, their capacity for video reasoning remains under-explored. We motivate framing video reasoning as t…

DescriptiveVideo Prediction