paper-with-me

홈 › Papers

BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and Editing

2023-05-24 · NeurIPS 2023 11 · Dongxu Li, Junnan Li, Steven C. H. Hoi

Subject-driven text-to-image generation models create novel renditions of an input subject based on text prompts. Existing models suffer from lengthy fine-tuning and difficulties preserving the subject fidelity. To overcome these limitations, we introduce BLIP-Diffusion, a new subject-driven image generation model that supports multimodal control which consumes inputs of subject images and text prompts. Unlike other subject-driven generation models, BLIP-Diffusion introduces a new multimodal encoder which is pre-trained to provide subject representation. We first pre-train the multimodal encoder following BLIP-2 to produce visual representation aligned with the text. Then we design a subject representation learning task which enables a diffusion model to leverage such visual representation and generates new subject renditions. Compared with previous methods such as DreamBooth, our model enables zero-shot subject-driven generation, and efficient fine-tuning for customized subject with up to 20x speedup. We also demonstrate that BLIP-Diffusion can be flexibly combined with existing techniques such as ControlNet and prompt-to-prompt to enable novel subject-driven generation and editing applications. Code and models will be released at https://github.com/salesforce/LAVIS/tree/main/projects/blip-diffusion. Project page at https://dxli94.github.io/BLIP-Diffusion-website/.

📄 PDF Abstract BibTeX arXiv:2305.14720

Code (1)

salesforce/lavis 공식 구현 pytorch

Tasks

Image GenerationPersonalized Image GenerationRepresentation LearningText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer

2024-10-19 · Ying Hu, Chenyi Zhuang, Pan Gao

Style transfer aims to fuse the artistic representation of a style image with the structural information of a content image. Existing methods train specific networks or utilize pre-trained models to learn content and sty…

Style Transfer

AnaDiffusion: Anatomically CompositionalLatent Diffusion for Controllable 3D Brain MRI Generation

2026-08-24 · Huiwen Han, Lulin Liu, Bangya Liu, Yuanhao Cai 외 arxiv

3D brain MRI generation has made significant advances in medical imaging, simulation, and controllable anatomical analysis. However, existing generative models typically synthesize 3D volumes monolithically, often overlo…

MedBLIP: Bootstrapping Language-Image Pre-training from 3D Medical Images and Texts

2023-05-18 · Qiuhui Chen, Xinyue Hu, ZiRui Wang, Yi Hong

Vision-language pre-training (VLP) models have been demonstrated to be effective in many computer vision applications. In this paper, we consider developing a VLP model in the medical domain for making computer-aided dia…

Medical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)+2

DiffuseKronA: A Parameter Efficient Fine-tuning Method for Personalized Diffusion Models

2024-02-27 · Shyam Marjit, Harshit Singh, Nityanand Mathur, Sayak Paul 외

In the realm of subject-driven text-to-image (T2I) generative models, recent developments like DreamBooth and BLIP-Diffusion have led to impressive results yet encounter limitations due to their intensive fine-tuning dem…

Image Generationparameter-efficient fine-tuningSensitivity

Susceptibility Distortion Correction of Diffusion MRI with a single Phase-Encoding Direction

2025-08-18 · Sedigheh Dargahi, Sylvain Bouix, Christian Desrosiers arxiv

Diffusion MRI (dMRI) is a valuable tool to map brain microstructure and connectivity by analyzing water molecule diffusion in tissue. However, acquiring dMRI data requires to capture multiple 3D brain volumes in a short …