paper-with-me

홈 › Papers

Adapting Image-to-Video Diffusion Models for Large-Motion Frame Interpolation

2024-12-22 · Luoxu Jin, Hiroshi Watanabe

With the development of video generation models has advanced significantly in recent years, we adopt large-scale image-to-video diffusion models for video frame interpolation. We present a conditional encoder designed to adapt an image-to-video model for large-motion frame interpolation. To enhance performance, we integrate a dual-branch feature extractor and propose a cross-frame attention mechanism that effectively captures both spatial and temporal information, enabling accurate interpolations of intermediate frames. Our approach demonstrates superior performance on the Fr\'echet Video Distance (FVD) metric when evaluated against other state-of-the-art approaches, particularly in handling large motion scenarios, highlighting advancements in generative-based methodologies.

📄 PDF Abstract BibTeX arXiv:2412.17042

Code (0)

등록된 구현이 없습니다.

Tasks

Video Frame InterpolationVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation

2024-08-27 · Xiaojuan Wang, Boyang Zhou, Brian Curless, Ira Kemelmacher-Shlizerman 외

We present a method for generating video sequences with coherent motion between a pair of input key frames. We adapt a pretrained large-scale image-to-video diffusion model (originally trained to generate videos moving f…

AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction

2024-06-10 · Zhen Xing, Qi Dai, Zejia Weng, Zuxuan Wu 외

Text-guided video prediction (TVP) involves predicting the motion of future frames from the initial frame according to an instruction, which has wide applications in virtual reality, robotics, and content creation. Previ…

Language ModellingLarge Language ModelVideo Prediction

MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation

2025-09-30 · Chenhui Zhu, Yilu Wu, Shuai Wang, Gangshan Wu 외 arxiv

Image-to-video generation has made remarkable progress with the advancements in diffusion models, yet generating videos with realistic motion remains highly challenging. This difficulty arises from the complexity of accu…

Zero-shot GeneralizationVideo Generation

MAD: Motion Appearance Decoupling for efficient Driving World Models

2026-01-14 · Ahmad Rahimi, Valentin Gerard, Eloi Zablocki, Matthieu Cord 외 arxiv

Recent video diffusion models generate photorealistic, temporally coherent videos, yet they fall short as reliable world models for autonomous driving, where structured motion and physically consistent interactions are e…

Autonomous Driving

IE2Video: Adapting Pretrained Diffusion Models for Event-Based Video Reconstruction

2025-12-04 · Dmitrii Torbunov, Onur Okuducu, Yi Huang, Odera Dim 외 arxiv

Continuous video monitoring in surveillance, robotics, and wearable systems faces a fundamental power constraint: conventional RGB cameras consume substantial energy through fixed-rate capture. Event cameras offer sparse…

Event-Based Video Reconstruction