paper-with-me

홈 › Papers

TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation

2025-07-07 · Zonglin Lyu, Chen Chen

Video Frame Interpolation (VFI) aims to predict the intermediate frame $I_n$ (we use n to denote time in videos to avoid notation overload with the timestep $t$ in diffusion models) based on two consecutive neighboring frames $I_0$ and $I_1$. Recent approaches apply diffusion models (both image-based and video-based) in this task and achieve strong performance. However, image-based diffusion models are unable to extract temporal information and are relatively inefficient compared to non-diffusion methods. Video-based diffusion models can extract temporal information, but they are too large in terms of training scale, model size, and inference time. To mitigate the above issues, we propose Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation (TLB-VFI), an efficient video-based diffusion model. By extracting rich temporal information from video inputs through our proposed 3D-wavelet gating and temporal-aware autoencoder, our method achieves 20% improvement in FID on the most challenging datasets over recent SOTA of image-based diffusion models. Meanwhile, due to the existence of rich temporal information, our method achieves strong performance while having 3times fewer parameters. Such a parameter reduction results in 2.3x speed up. By incorporating optical flow guidance, our method requires 9000x less training data and achieves over 20x fewer parameters than video-based diffusion models. Codes and results are available at our project page: https://zonglinl.github.io/tlbvfi_page.

📄 PDF Abstract BibTeX arXiv:2507.04984

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Flow EstimationVideo Frame Interpolation

Similar Papers 제목 키워드 기반

Frame Interpolation with Consecutive Brownian Bridge Diffusion

2024-05-09 · Zonglin Lyu, Ming Li, Jianbo Jiao, Chen Chen

Recent work in Video Frame Interpolation (VFI) tries to formulate VFI as a diffusion-based conditional image generation problem, synthesizing the intermediate frame given a random noise and neighboring frames. Due to the…

Conditional Image GenerationImage GenerationVideo Frame Interpolation

Improving the Plausibility of Pressure Distributions Synthesized from Depth Image through Generative Modeling

2025-12-15 · Neevkumar Manavar, Hanno Gerd Meyer, Joachim Waßmuth, Barbara Hammer 외 arxiv

Monitoring contact pressure in hospital beds is essential for preventing pressure ulcers and enabling real-time patient assessment. Current methods can predict pressure maps but often lack physical plausibility, limiting…

BBDM: Image-to-image Translation with Brownian Bridge Diffusion Models

2022-05-16 · CVPR 2023 1 · Bo Li, Kaitao Xue, Bin Liu, Yu-Kun Lai

Image-to-image translation is an important and challenging problem in computer vision and image processing. Diffusion models (DM) have shown great potentials for high-quality image synthesis, and have gained competitive …

Image GenerationImage-to-Image TranslationTranslation

Fractional Diffusion Bridge Models

2025-11-03 · Gabriel Nobis, Maximilian Springenberg, Arina Belova, Rembert Daems 외 arxiv

We present Fractional Diffusion Bridge Models (FDBM), a novel generative diffusion bridge framework driven by an approximation of the rich and non-Markovian fractional Brownian motion (fBM). Real stochastic processes exh…

Protein Structure Prediction

Dialogue Planning via Brownian Bridge Stochastic Process for Goal-directed Proactive Dialogue

2023-05-09 · Jian Wang, Dongding Lin, Wenjie Li

Goal-directed dialogue systems aim to proactively reach a pre-determined target through multi-turn conversations. The key to achieving this task lies in planning dialogue paths that smoothly and coherently direct convers…

Dialogue Generation