paper-with-me

홈 › Papers

DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion Transformer

2026-05-15 · Qingji Dong, Hang Dong, Mingqin Chen, Rui Zhang, Yitong Wang arxiv

Large-scale pre-trained diffusion models have been extensively adopted for real-world image Super-Resolution because of their powerful generative priors through textual guidance. However, when super-resolving high-resolution images with patch-wise inference strategy, most existing diffusion-based SR methods tend to suffer from over-generation, due to the misalignment between the global prompt from LR image and the incomplete semantic information of local patches during each inference step. On the other hand, most existing methods also failed to generate detailed texture in local patches due to the overemphasis on global generation capabilities in network designs and training strategies. To address this issue, we present DreamSR, a novel SR model that suppresses local over-generation and improves fine-detail synthesis, thereby achieving visually faithful results with ultra-high-quality details. Specifically, we propose a dual-branch MM-ControlNet, where the ControlNet generates local textual feature with patch-level prompts while the pre-trained DiT provides global textual feature with global prompts, thereby mitigating over-generation and ensuring semantic consistency across patches. We also design a comprehensive training strategy with stage-specific data processing pipelines and a Receptive-Field Enhancement strategy, enhancing the model's capability to capture patch information and effectively restore local textures. Extensive experiments demonstrate that DreamSR outperforms state-of-the-art methods, providing high-quality SR results. Code and model are available at https://github.com/jerrydong0219/DreamSR.

📄 PDF Abstract BibTeX arXiv:2605.15682

Code (0)

등록된 구현이 없습니다.

Tasks

Image Super-Resolution

Similar Papers 제목 키워드 기반

Perception Consistency Ultrasound Image Super-resolution via Self-supervised CycleGAN

2020-12-28 · Heng Liu, Jianyong Liu, Tao Tao, Shudong Hou 외

Due to the limitations of sensors, the transmission medium and the intrinsic properties of ultrasound, the quality of ultrasound imaging is always not ideal, especially its low spatial resolution. To remedy this situatio…

Generative Adversarial NetworkImage EnhancementImage Super-ResolutionSSIM+1

UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention

2025-10-18 · Yuyao Zhang, Yu-Wing Tai arxiv

Ultra-high-resolution text-to-image generation is increasingly vital for applications requiring fine-grained textures and global structural fidelity, yet state-of-the-art text-to-image diffusion models such as FLUX and S…

Text-to-Image GenerationComputational Efficiency

Semantic Encoder Guided Generative Adversarial Face Ultra-Resolution Network

2022-11-18 · Xiang Wang, Yimin Yang, Qixiang Pang, Xiao Lu 외

Face super-resolution is a domain-specific image super-resolution, which aims to generate High-Resolution (HR) face images from their Low-Resolution (LR) counterparts. In this paper, we propose a novel face super-resolut…

Image Super-ResolutionSuper-Resolution

Deep Learning-based Synthetic High-Resolution In-Depth Imaging Using an Attachable Dual-element Endoscopic Ultrasound Probe

2023-09-13 · Hah Min Lew, Jae Seong Kim, Moon Hwan Lee, Jaegeun Park 외

Endoscopic ultrasound (EUS) imaging has a trade-off between resolution and penetration depth. By considering the in-vivo characteristics of human organs, it is necessary to provide clinicians with appropriate hardware sp…

Deep LearningImage-to-Image TranslationSuper-Resolution

Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models

2025-03-24 · CVPR 2025 1 · Jinjin Zhang, Qiuyu Huang, Junjie Liu, Xiefan Guo 외

In this paper, we present Diffusion-4K, a novel framework for direct ultra-high-resolution image synthesis using text-to-image diffusion models. The core advancements include: (1) Aesthetic-4K Benchmark: addressing the a…

4kImage Generation