paper-with-me

Papers

Video Dynamics Prior: An Internal Learning Approach for Robust Video Enhancements

2023-12-13 · NeurIPS 2023 11 · Gaurav Shrivastava, Ser-Nam Lim, Abhinav Shrivastava

In this paper, we present a novel robust framework for low-level vision tasks, including denoising, object removal, frame interpolation, and super-resolution, that does not require any external training data corpus. Our proposed approach directly learns the weights of neural modules by optimizing over the corrupted test sequence, leveraging the spatio-temporal coherence and internal statistics of videos. Furthermore, we introduce a novel spatial pyramid loss that leverages the property of spatio-temporal patch recurrence in a video across the different scales of the video. This loss enhances robustness to unstructured noise in both the spatial and temporal domains. This further results in our framework being highly robust to degradation in input frames and yields state-of-the-art results on downstream tasks such as denoising, object removal, and frame interpolation. To validate the effectiveness of our approach, we conduct qualitative and quantitative evaluations on standard video datasets such as DAVIS, UCF-101, and VIMEO90K-T.

📄 PDF Abstract BibTeX arXiv:2312.07835

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSuper-Resolution

Similar Papers 제목 키워드 기반

UnDIVE: Generalized Underwater Video Enhancement Using Generative Priors

2024-11-08 · Suhas Srinath, Aditya Chandrasekar, Hemang Jamadagni, Rajiv Soundararajan 외

With the rise of marine exploration, underwater imaging has gained significant attention as a research topic. Underwater video enhancement has become crucial for real-time computer vision tasks in marine exploration. How…

DenoisingDescriptiveVideo Enhancement

Object Aware Egocentric Online Action Detection

2024-06-03 · Joungbin An, Yunsu Park, Hyolim Kang, Seon Joo Kim

Advancements in egocentric video datasets like Ego4D, EPIC-Kitchens, and Ego-Exo4D have enriched the study of first-person human interactions, which is crucial for applications in augmented reality and assisted living. D…

Action DetectionObjectOnline Action DetectionScene Understanding

Motion Marionette: Rethinking Rigid Motion Transfer via Prior Guidance

2025-11-25 · Haoxuan Wang, Jiachen Tao, Junyi Wu, Gaowen Liu 외 arxiv

We present Motion Marionette, a zero-shot framework for rigid motion transfer from monocular source videos to single-view target images. Previous works typically employ geometric, generative, or simulation priors to guid…

Video Generation

Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning

2025-12-10 · Xinyu Liu, Hangjie Yuan, Yujie Wei, Jiazheng Xing 외 arxiv

Unified video models exhibit strong capabilities in understanding and generation, yet they struggle with reason-informed visual editing even when equipped with powerful internal vision-language models (VLMs). We attribut…

Video Generation

KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation

2026-08-17 · Jianjie Luo, Yiming Zhong, Haoming Shen, Yupeng Xiao 외 arxiv

Identity-preserving video generation (IPVG) requires synthesizing videos that are faithful to both reference subjects and text prompts. Existing methods are often hindered by high tuning costs or limited input-level enha…

Video Generation