paper-with-me

Papers

Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction

2024-07-15 · Lin Zhu, Yunlong Zheng, Yijun Zhang, Xiao Wang, Lizhi Wang, Hua Huang

Event-based video reconstruction has garnered increasing attention due to its advantages, such as high dynamic range and rapid motion capture capabilities. However, current methods often prioritize the extraction of temporal information from continuous event flow, leading to an overemphasis on low-frequency texture features in the scene, resulting in over-smoothing and blurry artifacts. Addressing this challenge necessitates the integration of conditional information, encompassing temporal features, low-frequency texture, and high-frequency events, to guide the Denoising Diffusion Probabilistic Model (DDPM) in producing accurate and natural outputs. To tackle this issue, we introduce a novel approach, the Temporal Residual Guided Diffusion Framework, which effectively leverages both temporal and frequency-based event priors. Our framework incorporates three key conditioning modules: a pre-trained low-frequency intensity estimation module, a temporal recurrent encoder module, and an attention-based high-frequency prior enhancement module. In order to capture temporal scene variations from the events at the current moment, we employ a temporal-domain residual image as the target for the diffusion model. Through the combination of these three conditioning paths and the temporal residual framework, our framework excels in reconstructing high-quality videos from event flow, mitigating issues such as artifacts and over-smoothing commonly observed in previous approaches. Extensive experiments conducted on multiple benchmark datasets validate the superior performance of our framework compared to prior event-based reconstruction methods.

📄 PDF Abstract BibTeX arXiv:2407.10636

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingEvent-Based Video ReconstructionVideo Reconstruction

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DESSERT: Diffusion-based Event-driven Single-frame Synthesis via Residual Training

2025-12-19 · Jiyun Kong, Jun-Hyuk Kim, Jong-Seok Lee arxiv

Video frame prediction extrapolates future frames from previous frames, but suffers from prediction errors in dynamic scenes due to the lack of information about the next frame. Event cameras address this limitation by c…

Video Frame InterpolationEvent-based Optical Flow

FADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face Restoration

2026-07-07 · Jin Jiang, Jia Wang, Panwen Hu, Weiran Zhao 외 arxiv

Video face restoration (VFR) aims to recover high-quality and temporally consistent facial details from severely degraded video sequences; however, existing methods still struggle to balance spatial fidelity and temporal…

GLIDE: Graph-guided Leap Inference for Diffusion Estimation of Spatio-Temporal Point Processes

2026-05-31 · Guanyu Zhou, Yao Liu, Yanglei Gan, Yuxiang Cai 외 arxiv

Spatio-temporal point processes (STPPs) provide a principled framework for modeling asynchronous events in continuous time and space. Recent diffusion-based approaches offer a flexible alternative to deterministic predic…

Point Processes

EGVD: Event-Guided Video Diffusion Model for Physically Realistic Large-Motion Frame Interpolation

2025-03-26 · Ziran Zhang, Xiaohui Li, Yihao Liu, Yujin Wang 외

Video frame interpolation (VFI) in scenarios with large motion remains challenging due to motion ambiguity between frames. While event cameras can capture high temporal resolution motion information, existing event-based…

Video Frame Interpolation

Pose-Guided Residual Refinement for Interpretable Text-to-Motion Generation and Editing

2025-12-27 · Sukhyun Jeong, Yong-Hoon Choi arxiv

Text-based 3D motion generation aims to automatically synthesize diverse motions from natural-language descriptions to extend user creativity, whereas motion editing modifies an existing motion sequence in response to te…