paper-with-me

Papers

Taming Real-World Space-Time Video Super-Resolution with One-Step Diffusion

2026-01-28 · Shuoyan Wei, Feng Li, Chen Zhou, Runmin Cong, Yao Zhao, Huihui Bai arxiv

Diffusion models have demonstrated exceptional success in video super-resolution (VSR), exhibiting powerful capabilities for generating fine-grained details. However, their potential for space-time video super-resolution (STVSR), which necessitates not only recovering realistic high-resolution visual content but also improving the frame rate with coherent temporal dynamics, remains largely underexplored. Moreover, existing STVSR methods predominantly address spatiotemporal upsampling under simple degradation assumptions, thus failing in real-world scenarios with complex unknown degradations. To address these challenges, we propose OSDEnhancer, the first framework that achieves robust STVSR in one-step diffusion. OSDEnhancer begins with a linear initialization to establish essential spatiotemporal structures and adapt the model for one-step reconstruction. It then applies a divide-and-conquer strategy, introducing the temporal coherence (TC) and texture enrichment (TE) LoRAs that progressively specialize in inter-frame dynamics modeling and fine-grained texture recovery, respectively, while collaborating during inference for enhanced overall performance. A bidirectional VAE decoder employs deformable recurrent blocks to leverage the multi-scale structure of the vanilla VAE, enhancing latent-to-pixel reconstruction through joint multi-scale deformable aggregation and inter-frame feature propagation. Experimental results demonstrate that the proposed method attains state-of-the-art performance with superior generalization in real-world scenarios. The code is available at https://github.com/W-Shuoyan/OSDEnhancer.

📄 PDF Abstract BibTeX arXiv:2601.20308

Code (0)

등록된 구현이 없습니다.

Tasks

Space-time Video Super-resolution

Similar Papers 제목 키워드 기반

ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling

2025-12-01 · Qisen Wang, Yifan Zhao, Peisen Shen, Jialu Li 외 arxiv

Although prevailing camera-controlled video generation models can produce cinematic results, lifting them directly to the generation of 3D-consistent and high-fidelity time-synchronized multi-view videos remains challeng…

Data AugmentationVideo Generation

HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation

2025-04-30 · Haiyang Zhou, Wangbo Yu, Jiawen Guan, Xinhua Cheng 외

The rapid advancement of diffusion models holds the promise of revolutionizing the application of VR and AR technologies, which typically require scene-level 4D assets for user experience. Nonetheless, existing diffusion…

Depth EstimationScene GenerationVideo Generation

TITAN-Guide: Taming Inference-Time AligNment for Guided Text-to-Video Diffusion Models

2025-08-01 · Christian Simon, Masato Ishii, Akio Hayakawa, Zhi Zhong 외 arxiv

In the recent development of conditional diffusion models still require heavy supervised fine-tuning for performing control on a category of tasks. Training-free conditioning via guidance with off-the-shelf models is a f…

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation

2025-10-10 · Chuanrui Zhang, Zhengxian Wu, Guanxing Lu, Yansong Tang 외 arxiv

Learned world models hold significant potential as neural simulators for robotic manipulation. However, prevalent 2D video-based models inherently lack the spatial and kinematic reasoning crucial for physical interaction…

Reinforcement Learning

OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder

2026-03-17 · Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu 외 arxiv

Existing diffusion-based 3D scene generation methods primarily operate in 2D image/video latent spaces, which makes maintaining cross-view appearance and geometric consistency inherently challenging. To bridge this gap, …

Scene Generation