paper-with-me

홈 › Papers

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

2025-04-10 · Zhong-Yu Li, Ruoyi Du, Juncheng Yan, Le Zhuo, Zhen Li, Peng Gao, Zhanyu Ma, Ming-Ming Cheng

Recent progress in diffusion models significantly advances various image generation tasks. However, the current mainstream approach remains focused on building task-specific models, which have limited efficiency when supporting a wide range of different needs. While universal models attempt to address this limitation, they face critical challenges, including generalizable task instruction, appropriate task distributions, and unified architectural design. To tackle these challenges, we propose VisualCloze, a universal image generation framework, which supports a wide range of in-domain tasks, generalization to unseen ones, unseen unification of multiple tasks, and reverse generation. Unlike existing methods that rely on language-based task instruction, leading to task ambiguity and weak generalization, we integrate visual in-context learning, allowing models to identify tasks from visual demonstrations. Meanwhile, the inherent sparsity of visual task distributions hampers the learning of transferable knowledge across tasks. To this end, we introduce Graph200K, a graph-structured dataset that establishes various interrelated tasks, enhancing task density and transferable knowledge. Furthermore, we uncover that our unified image generation formulation shared a consistent objective with image infilling, enabling us to leverage the strong generative priors of pre-trained infilling models without modifying the architectures.

📄 PDF Abstract BibTeX arXiv:2504.07960

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationIn-Context Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models

2025-09-26 · Komal Kumar, Rao Muhammad Anwer, Fahad Shahbaz Khan, Salman Khan 외 arxiv

Efficient fine-tuning of pre-trained Text-to-Image (T2I) models involves adjusting the model to suit a particular task or dataset while minimizing computational resources and limiting the number of trainable parameters. …

Image Generation

Visual Bridge: Universal Visual Perception Representations Generating

2025-11-11 · Yilin Gao, Shuguang Dou, Junzhou Li, Zhiheng Yu 외 arxiv

Recent advances in diffusion models have achieved remarkable success in isolated computer vision tasks such as text-to-image generation, depth estimation, and optical flow. However, these models are often restricted by a…

Text-to-Image GenerationDomain GeneralizationDepth EstimationText Retrieval

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

2024-12-10 · CVPR 2025 1 · Xi Chen, Zhifei Zhang, He Zhang, Yuqian Zhou 외

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs an…

Image GenerationVideo Generation

Generative Universal Verifier as Multimodal Meta-Reasoner

2025-10-15 · Xinchen Zhang, Xiaoying Zhang, Youbin Wu, Yanbin Cao 외 arxiv

We introduce Generative Universal Verifier, a novel concept and plugin designed for next-generation multimodal reasoning in vision-language models and unified multimodal models, providing the fundamental capability of re…

Multimodal ReasoningImage Generation

MoE-DiffIR: Task-customized Diffusion Priors for Universal Compressed Image Restoration

2024-07-15 · Yulin Ren, Xin Li, Bingchen Li, Xingrui Wang 외

We present MoE-DiffIR, an innovative universal compressed image restoration (CIR) method with task-customized diffusion priors. This intends to handle two pivotal challenges in the existing CIR methods: (i) lacking adapt…

Image RestorationMixture-of-ExpertsTexture Synthesis