paper-with-me

홈 › Papers

FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion

2024-12-12 · Haonan Qiu, Shiwei Zhang, Yujie Wei, Ruihang Chu, Hangjie Yuan, Xiang Wang, Yingya Zhang, Ziwei Liu

Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack of high-resolution data and constrained computation resources, hampering their ability to generate high-fidelity images or videos at higher resolutions. Recent efforts have explored tuning-free strategies to exhibit the untapped potential higher-resolution visual generation of pre-trained models. However, these methods are still prone to producing low-quality visual content with repetitive patterns. The key obstacle lies in the inevitable increase in high-frequency information when the model generates visual content exceeding its training resolution, leading to undesirable repetitive patterns deriving from the accumulated errors. To tackle this challenge, we propose FreeScale, a tuning-free inference paradigm to enable higher-resolution visual generation via scale fusion. Specifically, FreeScale processes information from different receptive scales and then fuses it by extracting desired frequency components. Extensive experiments validate the superiority of our paradigm in extending the capabilities of higher-resolution visual generation for both image and video models. Notably, compared with the previous best-performing method, FreeScale unlocks the generation of 8k-resolution images for the first time.

📄 PDF Abstract BibTeX arXiv:2412.09626

Code (0)

등록된 구현이 없습니다.

Tasks

8k

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation

2026-04-12 · Chenhan Jiang, Yu Chen, Qingwen Zhang, Jifei Song 외 arxiv

The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photo…

Novel View Synthesis

FreeScale: Distributed Training for Sequence Recommendation Models with Minimal Scaling Cost

2026-04-27 · Chenhao Feng, Haoli Zhang, Shakhzod Ali-Zade, Yanli Zhao 외 arxiv

Modern industrial Deep Learning Recommendation Models typically extract user preferences through the analysis of sequential interaction histories, subsequently generating predictions based on these derived interests. The…

FaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolution

2024-11-27 · CVPR 2025 1 · Junyang Chen, Jinshan Pan, Jiangxin Dong

Faithful image super-resolution (SR) not only needs to recover images that appear realistic, similar to image generation tasks, but also requires that the restored images maintain fidelity and structural consistency with…

Image GenerationImage Super-ResolutionSuper-Resolution

EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model

2026-04-11 · Kunho Kim, Sumin Seo, Yongjun Cho, Hyungjin Chung arxiv

We propose EditCrafter, a high-resolution image editing method that operates without tuning, leveraging pretrained text-to-image (T2I) diffusion models to process images at resolutions significantly exceeding those used …

Image Editing

Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation

2024-02-16 · Lanqing Guo, Yingqing He, Haoxin Chen, Menghan Xia 외

Diffusion models have proven to be highly effective in image and video generation; however, they still face composition challenges when generating images of varying sizes due to single-scale training data. Adapting large…

Video Generation