paper-with-me

Papers

One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution

2025-11-21 · Yushun Fang, Yuxiang Chen, Shibo Yin, Qiang Hu, Jiangchao Yao, Ya Zhang, Xiaoyun Zhang, Yanfeng Wang arxiv

Recent advances in diffusion-based real-world image super-resolution (Real-ISR) have demonstrated remarkable perceptual quality, yet the balance between fidelity and controllability remains a problem: multi-step diffusion-based methods suffer from generative diversity and randomness, resulting in low fidelity, while one-step methods lose control flexibility due to fidelity-specific finetuning. In this paper, we present ODTSR, a one-step diffusion transformer based on Qwen-Image that performs Real-ISR considering fidelity and controllability simultaneously: a newly introduced visual stream receives low-quality images (LQ) with adjustable noise (Control Noise), and the original visual stream receives LQs with consistent noise (Prior Noise), forming the Noise-hybrid Visual Stream (NVS) design. ODTSR further employs Fidelity-aware Adversarial Training (FAA) to enhance controllability and achieve one-step inference. Extensive experiments demonstrate that ODTSR not only achieves state-of-the-art (SOTA) performance on generic Real-ISR, but also enables prompt controllability on challenging scenarios such as real-world scene text image super-resolution (STISR) of Chinese characters without training on specific datasets. Codes are available at https://github.com/RedMediaTech/ODTSR.

📄 PDF Abstract BibTeX arXiv:2511.17138

Code (0)

등록된 구현이 없습니다.

Tasks

Image Super-Resolution

Similar Papers 제목 키워드 기반

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

2026-05-28 · Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou 외 arxiv

Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains challenging. Interactive world models re…

Video Generation

Time-Aware One Step Diffusion Network for Real-World Image Super-Resolution

2025-08-22 · Tianyi Zhang, Zheng-Peng Duan, Peng-Tao Jiang, Bo Li 외 arxiv

Diffusion-based real-world image super-resolution (Real-ISR) methods have demonstrated impressive performance.To achieve efficient Real-ISR, many works employ Variational Score Distillation (VSD) to distill pre-trained s…

Image Super-Resolution

Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

2025-08-30 · Xuechao Zou, Shun Zhang, Xing Fu, Yue Li 외 arxiv

Controllable face generation poses critical challenges in generative modeling due to the intricate balance required between semantic controllability and photorealism. While existing approaches struggle with disentangling…

Zero-shot Generalization

Controllable Shadow Generation with Single-Step Diffusion Models from Synthetic Data

2024-12-16 · Onur Tasar, Clément Chadebec, Benjamin Aubin

Realistic shadow generation is a critical component for high-quality image compositing and visual effects, yet existing methods suffer from certain limitations: Physics-based approaches require a 3D scene geometry, which…

JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on

2025-08-25 · Aowen Wang, Wei Li, Hao Luo, Mengxing Ao 외 arxiv

Virtual try-on systems have long been hindered by heavy reliance on human body masks, limited fine-grained control over garment attributes, and poor generalization to real-world, in-the-wild scenarios. In this paper, we …

Image GenerationVirtual Try-on