paper-with-me

홈 › Papers

PromptFix: You Prompt and We Fix the Photo

2024-05-27 · Yongsheng Yu, Ziyun Zeng, Hang Hua, Jianlong Fu, Jiebo Luo

Diffusion models equipped with language models demonstrate excellent controllability in image generation tasks, allowing image processing to adhere to human instructions. However, the lack of diverse instruction-following data hampers the development of models that effectively recognize and execute user-customized instructions, particularly in low-level tasks. Moreover, the stochastic nature of the diffusion process leads to deficiencies in image generation or editing tasks that require the detailed preservation of the generated images. To address these limitations, we propose PromptFix, a comprehensive framework that enables diffusion models to follow human instructions to perform a wide variety of image-processing tasks. First, we construct a large-scale instruction-following dataset that covers comprehensive image-processing tasks, including low-level tasks, image editing, and object creation. Next, we propose a high-frequency guidance sampling method to explicitly control the denoising process and preserve high-frequency details in unprocessed areas. Finally, we design an auxiliary prompting adapter, utilizing Vision-Language Models (VLMs) to enhance text prompts and improve the model's task generalization. Experimental results show that PromptFix outperforms previous methods in various image-processing tasks. Our proposed model also achieves comparable inference efficiency with these baseline models and exhibits superior zero-shot capabilities in blind restoration and combination tasks. The dataset and code are available at https://www.yongshengyu.com/PromptFix-Page.

📄 PDF Abstract BibTeX arXiv:2405.16785

Code (1)

yeates/promptfix 공식 구현 pytorch

Tasks

DenoisingImage GenerationInstruction Following

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

PromptFix: Few-shot Backdoor Removal via Adversarial Prompt Tuning

2024-06-06 · Tianrong Zhang, Zhaohan Xi, Ting Wang, Prasenjit Mitra 외

Pre-trained language models (PLMs) have attracted enormous attention over the past few years with their unparalleled performances. Meanwhile, the soaring cost to train PLMs as well as their amazing generalizability have …

A simulation study to distinguish prompt photon from $π^0$ and beam halo in a granular calorimeter using deep networks

2018-08-12 · Shamik Ghosh, Abhirami Harilal, A. R. Sahasransu, Ritesh Kumar Singh 외

In a hadron collider environment identification of prompt photons originating in a hard partonic scattering process and rejection of non-prompt photons coming from hadronic jets or from beam related sources, is the first…

Personalized Image Filter: Mastering Your Photographic Style

2025-10-19 · Chengxuan Zhu, Shuchen Weng, Jiacong Fang, Peixuan Zhang 외 arxiv

Photographic style, as a composition of certain photographic concepts, is the charm behind renowned photographers. But learning and transferring photographic style need a profound understanding of how the photo is edited…

Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning

2024-07-05 · Mainak Singha, Ankit Jha, Divyam Gupta, Pranav Singla 외

We address the challenges inherent in sketch-based image retrieval (SBIR) across various settings, including zero-shot SBIR, generalized zero-shot SBIR, and fine-grained zero-shot SBIR, by leveraging the vision-language …

AllImage RetrievalPrompt LearningRetrieval+2

AnyPhoto: Multi-Person Identity Preserving Image Generation with ID Adaptive Modulation on Location Canvas

2026-03-16 · Longhui Yuan arxiv

Multi-person identity-preserving generation requires binding multiple reference faces to specified locations under a text prompt. Strong identity/layout conditions often trigger copy-paste shortcuts and weaken prompt-dri…

Image Generation