paper-with-me

Papers

EventDiff: A Unified and Efficient Diffusion Model Framework for Event-based Video Frame Interpolation

2025-05-13 · Hanle Zheng, Xujie Han, Zegang Peng, Shangbin Zhang, Guangxun Du, Zhuo Zou, XiLin Wang, Jibin Wu, Hao Guo, Lei Deng

Video Frame Interpolation (VFI) is a fundamental yet challenging task in computer vision, particularly under conditions involving large motion, occlusion, and lighting variation. Recent advancements in event cameras have opened up new opportunities for addressing these challenges. While existing event-based VFI methods have succeeded in recovering large and complex motions by leveraging handcrafted intermediate representations such as optical flow, these designs often compromise high-fidelity image reconstruction under subtle motion scenarios due to their reliance on explicit motion modeling. Meanwhile, diffusion models provide a promising alternative for VFI by reconstructing frames through a denoising process, eliminating the need for explicit motion estimation or warping operations. In this work, we propose EventDiff, a unified and efficient event-based diffusion model framework for VFI. EventDiff features a novel Event-Frame Hybrid AutoEncoder (HAE) equipped with a lightweight Spatial-Temporal Cross Attention (STCA) module that effectively fuses dynamic event streams with static frames. Unlike previous event-based VFI methods, EventDiff performs interpolation directly in the latent space via a denoising diffusion process, making it more robust across diverse and challenging VFI scenarios. Through a two-stage training strategy that first pretrains the HAE and then jointly optimizes it with the diffusion model, our method achieves state-of-the-art performance across multiple synthetic and real-world event VFI datasets. The proposed method outperforms existing state-of-the-art event-based VFI methods by up to 1.98dB in PSNR on Vimeo90K-Triplet and shows superior performance in SNU-FILM tasks with multiple difficulty levels. Compared to the emerging diffusion-based VFI approach, our method achieves up to 5.72dB PSNR gain on Vimeo90K-Triplet and 4.24X faster inference.

📄 PDF Abstract BibTeX arXiv:2505.08235

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingImage ReconstructionMotion EstimationOptical Flow EstimationTripletVideo Frame Interpolation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models

2026-02-22 · Gang Xu, Zhiyu Zhu, Junhui Hou arxiv

Event cameras excel at high-speed, low-power, and high-dynamic-range scene perception. However, as they fundamentally record only relative intensity changes rather than absolute intensity, the resulting data streams suff…

Video Frame Interpolation

GenRec: Unifying Video Generation and Recognition with Diffusion Models

2024-08-27 · Zejia Weng, Xitong Yang, Zhen Xing, Zuxuan Wu 외

Video diffusion models are able to generate high-quality videos by learning strong spatial-temporal priors on large-scale datasets. In this paper, we aim to investigate whether such priors derived from a generative proce…

Image to Video GenerationVideo GenerationVideo Recognition

Omni-Video: Democratizing Unified Video Understanding and Generation

2025-07-08 · Zhiyu Tan, Hao Yang, Luozheng Qin, Jia Gong 외

Notable breakthroughs in unified understanding and generation modeling have led to remarkable advancements in image understanding, reasoning, production and editing, yet current foundational models predominantly focus on…

Video GenerationVideo Understanding

Bridging Video Understanding and Generation in a Unified Framework

2026-06-30 · Yuqi Wang, Runyi Li, Ruoyu Feng, Renjie Chen 외 arxiv

Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexplored. A central challenge is that video…

Video GenerationImage Generation

ER3: A Unified Framework for Event Retrieval, Recognition and Recounting

2017-07-01 · CVPR 2017 7 · Zhanning Gao, Gang Hua, Dong-Qing Zhang, Nebojsa Jojic 외

We develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alig…

Language ModelingLanguage ModellingRetrieval