paper-with-me

홈 › Papers

In-Context Learning Unlocked for Diffusion Models

2023-05-01 · NeurIPS 2023 11 · Zhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen, Pengcheng He, Weizhu Chen, Zhangyang Wang, Mingyuan Zhou

We present Prompt Diffusion, a framework for enabling in-context learning in diffusion-based generative models. Given a pair of task-specific example images, such as depth from/to image and scribble from/to image, and a text guidance, our model automatically understands the underlying task and performs the same task on a new query image following the text guidance. To achieve this, we propose a vision-language prompt that can model a wide range of vision-language tasks and a diffusion model that takes it as input. The diffusion model is trained jointly over six different tasks using these prompts. The resulting Prompt Diffusion model is the first diffusion-based vision-language foundation model capable of in-context learning. It demonstrates high-quality in-context generation on the trained tasks and generalizes effectively to new, unseen vision tasks with their respective prompts. Our model also shows compelling text-guided image editing results. Our framework aims to facilitate research into in-context learning for computer vision. We share our code and pre-trained models at https://github.com/Zhendong-Wang/Prompt-Diffusion.

📄 PDF Abstract BibTeX arXiv:2305.01115

Code (1)

zhendong-wang/prompt-diffusion 공식 구현 pytorch

Tasks

In-Context Learningtext-guided-image-editing

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

ModelLock: Locking Your Model With a Spell

2024-05-25 · Yifeng Gao, Yuhua Sun, Xingjun Ma, Zuxuan Wu 외

This paper presents a novel model protection paradigm ModelLock that locks (destroys) the performance of a model on normal clean data so as to make it unusable or unextractable without the right key. Specifically, we pro…

image-classificationImage Classificationmodeltext-guided-image-editing

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

2024-12-05 · Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan 외

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency…

Portrait AnimationVideo Generation

A Survey on Diffusion Models for Inverse Problems

2024-09-30 · Giannis Daras, Hyungjin Chung, Chieh-Hsin Lai, Yuki Mitsufuji 외

Diffusion models have become increasingly popular for generative modeling due to their ability to generate high-quality samples. This has unlocked exciting new possibilities for solving inverse problems, especially in im…

Image RestorationSurvey

Unlocking Prompt Infilling Capability for Diffusion Language Models

2026-04-04 · Yoshinari Fujinuma, Keisuke Sakaguchi arxiv

Masked diffusion language models (dLMs) generate text through bidirectional denoising, yet this capability remains locked for infilling prompts. This limitation is an artifact of the current supervised finetuning (SFT) c…

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers

2024-12-17 · Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi 외

Recent research arXiv:2410.15027 arXiv:2410.23775 has highlighted the inherent in-context generation capabilities of pretrained diffusion transformers (DiTs), enabling them to seamlessly adapt to diverse visual tasks wit…

ArticlesForm