paper-with-me

Papers

ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained Guidance

2024-06-24 · Shuwei Shi, Wenbo Li, Yuechen Zhang, Jingwen He, Biao Gong, Yinqiang Zheng

Diffusion models excel at producing high-quality images; however, scaling to higher resolutions, such as 4K, often results in over-smoothed content, structural distortions, and repetitive patterns. To this end, we introduce ResMaster, a novel, training-free method that empowers resolution-limited diffusion models to generate high-quality images beyond resolution restrictions. Specifically, ResMaster leverages a low-resolution reference image created by a pre-trained diffusion model to provide structural and fine-grained guidance for crafting high-resolution images on a patch-by-patch basis. To ensure a coherent global structure, ResMaster meticulously aligns the low-frequency components of high-resolution patches with the low-resolution reference at each denoising step. For fine-grained guidance, tailored image prompts based on the low-resolution reference and enriched textual prompts produced by a vision-language model are incorporated. This approach could significantly mitigate local pattern distortions and improve detail refinement. Extensive experiments validate that ResMaster sets a new benchmark for high-resolution image generation and demonstrates promising efficiency. The project page is https://shuweis.github.io/ResMaster .

📄 PDF Abstract BibTeX arXiv:2406.16476

Code (0)

등록된 구현이 없습니다.

Tasks

4kDenoisingImage GenerationLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

CogView: Mastering Text-to-Image Generation via Transformers

2021-05-26 · NeurIPS 2021 12 · Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 외

Text-to-Image generation in the general domain has long been an open problem, which requires both a powerful generative model and cross-modal understanding. We propose CogView, a 4-billion-parameter Transformer with VQ-V…

Image GenerationSuper-ResolutionText to Image GenerationText-to-Image Generation+1

DeepRemaster: Temporal Source-Reference Attention Networks for Comprehensive Video Enhancement

2020-09-18 · Satoshi Iizuka, Edgar Simo-Serra

The remastering of vintage film comprises of a diversity of sub-tasks including super-resolution, noise removal, and contrast enhancement which aim to restore the deteriorated film medium to its original state. Additiona…

ColorizationDiversitySuper-ResolutionVideo Enhancement

End-to-end Music Remastering System Using Self-supervised and Adversarial Training

2022-02-17 · Junghyun Koo, Seungryeol Paik, Kyogu Lee

Mastering is an essential step in music production, but it is also a challenging task that has to go through the hands of experienced audio engineers, where they adjust tone, space, and volume of a song. Remastering foll…

Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

2024-10-24 · Zhangheng Li, Keen You, Haotian Zhang, Di Feng 외

Building a generalist model for user interface (UI) understanding is challenging due to various foundational issues, such as platform diversity, resolution variation, and data limitation. In this paper, we introduce Ferr…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+2

Structural Regularities of Cinema SDR-to-HDR Mapping in a Controlled Mastering Workflow: A Pixel-wise Case Study on ASC StEM2

2026-04-07 · Xin Zhang, Xiaoyi Chen arxiv

We present an empirical case study of cinema SDR-to-HDR mapping using ASC StEM2, a rare common-source dataset containing EXR scene-referred images and matched SDR/HDR cinema release masters from the same ACES-based maste…