paper-with-me

홈 › Papers

Learning to Refocus with Video Diffusion Models

2025-12-22 · SaiKiran Tedla, Zhoutong Zhang, Xuaner Zhang, Shumian Xin arxiv

Focus is a cornerstone of photography, yet autofocus systems often fail to capture the intended subject, and users frequently wish to adjust focus after capture. We introduce a novel method for realistic post-capture refocusing using video diffusion models. From a single defocused image, our approach generates a perceptually accurate focal stack, represented as a video sequence, enabling interactive refocusing and unlocking a range of downstream applications. We release a large-scale focal stack dataset acquired under diverse real-world smartphone conditions to support this work and future research. Our method consistently outperforms existing approaches in both perceptual quality and robustness across challenging scenarios, paving the way for more advanced focus-editing capabilities in everyday photography. Code and data are available at www.learn2refocus.github.io

📄 PDF Abstract BibTeX arXiv:2512.19823

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Grounded Text-to-Image Synthesis with Attention Refocusing

2023-06-08 · CVPR 2024 1 · Quynh Phung, Songwei Ge, Jia-Bin Huang

Driven by the scalable diffusion models trained on large-scale datasets, text-to-image synthesis methods have shown compelling results. However, these models still fail to precisely follow the text prompt involving multi…

Image Generation

Re-Attentional Controllable Video Diffusion Editing

2024-12-16 · Yuanzhi Wang, Yong Li, Mengyi Liu, Xiaoya Zhang 외

Editing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. Recent studies have explored and exploite…

DenoisingVideo Editing

DiffCamera: Arbitrary Refocusing on Images

2025-09-30 · Yiyang Wang, Xi Chen, Xiaogang Xu, Yu Liu 외 arxiv

The depth-of-field (DoF) effect, which introduces aesthetically pleasing blur, enhances photographic quality but is fixed and difficult to modify once the image has been created. This becomes problematic when the applied…

Small Clips, Big Gains: Learning Long-Range Refocused Temporal Information for Video Super-Resolution

2025-05-04 · Xingyu Zhou, Wei Long, Jingbo Lu, Shiyin Jiang 외

Video super-resolution (VSR) can achieve better performance compared to single image super-resolution by additionally leveraging temporal information. In particular, the recurrent-based VSR model exploits long-range temp…

Computational EfficiencyImage Super-ResolutionSuper-ResolutionVideo Super-Resolution

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding

2025-06-02 · Hosu Lee, Junho Kim, Hyunjun Kim, Yong Man Ro

Recent progress in Large Multi-modal Models (LMMs) has enabled effective vision-language reasoning, yet the ability to understand video content remains constrained by suboptimal frame selection strategies. Existing appro…