paper-with-me

Papers

Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment

2025-05-24 · Bryan Sangwoo Kim, Jeongsol Kim, Jong Chul Ye

Modern single-image super-resolution (SISR) models deliver photo-realistic results at the scale factors on which they are trained, but collapse when asked to magnify far beyond that regime. We address this scalability bottleneck with Chain-of-Zoom (CoZ), a model-agnostic framework that factorizes SISR into an autoregressive chain of intermediate scale-states with multi-scale-aware prompts. CoZ repeatedly re-uses a backbone SR model, decomposing the conditional probability into tractable sub-problems to achieve extreme resolutions without additional training. Because visual cues diminish at high magnifications, we augment each zoom step with multi-scale-aware text prompts generated by a vision-language model (VLM). The prompt extractor itself is fine-tuned using Generalized Reward Policy Optimization (GRPO) with a critic VLM, aligning text guidance towards human preference. Experiments show that a standard 4x diffusion SR model wrapped in CoZ attains beyond 256x enlargement with high perceptual quality and fidelity. Project Page: https://bryanswkim.github.io/chain-of-zoom/ .

📄 PDF Abstract BibTeX arXiv:2505.18600

Code (0)

등록된 구현이 없습니다.

Tasks

Image Super-ResolutionLanguage ModelingLanguage ModellingSuper-Resolution

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance

2026-05-18 · Jiale Shi, Jiarui Hu, Zesong Yang, Kaixuan Luan 외 arxiv

We introduce GaussianZoom, a generative zoom-in 3D reconstruction system with an iterative progressive framework that combines geometry-consistent scene modeling and multi-scale semantic reasoning to enable high-fidelity…

3D Reconstruction

OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution

2026-09-06 · Shubhashis Roy Dipta, Sourajit Saha, Shaswati Saha, Nobin Sarwar hf

Recursive Super-Resolution (SR) extends fixed-scale SR to extreme magnification by repeatedly feeding predictions back into the same model, analogous to zooming an image repeatedly. However, ground truth availability at …

Generative Powers of Ten

2023-12-04 · CVPR 2024 1 · Xiaojuan Wang, Janne Kontkanen, Brian Curless, Steve Seitz 외

We present a method that uses a text-to-image model to generate consistent content across multiple image scales, enabling extreme semantic zooms into a scene, e.g., ranging from a wide-angle landscape view of a forest to…

Image Super-ResolutionSuper-Resolution

EarthGen: Generating the World from Top-Down Views

2024-09-02 · Ansh Sharma, Albert Xiao, Praneet Rathi, Rohit Kundu 외

In this work, we present a novel method for extensive multi-scale generative terrain modeling. At the core of our model is a cascade of superresolution diffusion models that can be combined to produce consistent images a…

Scene GenerationSuper-Resolution

Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance

2025-10-24 · Minxing Luo, Linlong Fan, Wang Qiushi, Ge Wu 외 arxiv

Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image quality and textual readability. To address this, we introduce TIGER (Text…

Image Super-ResolutionImage Enhancement