paper-with-me

Papers

LLM-guided Instance-level Image Manipulation with Diffusion U-Net Cross-Attention Maps

2025-01-23 · Andrey Palaev, Adil Khan, Syed M. Ahsan Kazmi

The advancement of text-to-image synthesis has introduced powerful generative models capable of creating realistic images from textual prompts. However, precise control over image attributes remains challenging, especially at the instance level. While existing methods offer some control through fine-tuning or auxiliary information, they often face limitations in flexibility and accuracy. To address these challenges, we propose a pipeline leveraging Large Language Models (LLMs), open-vocabulary detectors, cross-attention maps and intermediate activations of diffusion U-Net for instance-level image manipulation. Our method detects objects mentioned in the prompt and present in the generated image, enabling precise manipulation without extensive training or input masks. By incorporating cross-attention maps, our approach ensures coherence in manipulated images while controlling object positions. Our method enables precise manipulations at the instance level without fine-tuning or auxiliary information such as masks or bounding boxes. Code is available at https://github.com/Palandr123/DiffusionU-NetLLM

📄 PDF Abstract BibTeX arXiv:2501.14046

Code (1)

palandr123/diffusionu-netllm 공식 구현 pytorch

Tasks

Image GenerationImage Manipulation

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
U-Net 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation

2025-11-25 · Yuhan Wu, Tiantian Wei, Shuo Wang, ZhiChao Wang 외 arxiv

Interactive articulated manipulation requires long-horizon, multi-step interactions with appliances while maintaining physical consistency. Existing vision-language and diffusion-based policies struggle to generalize acr…

S$^2$-Diffusion: Generalizing from Instance-level to Category-level Skills in Robot Manipulation

2025-02-13 · Quantao Yang, Michael C. Welle, Danica Kragic, Olov Andersson

Recent advances in skill learning has propelled robot manipulation to new heights by enabling it to learn complex manipulation tasks from a practical number of demonstrations. However, these skills are often limited to t…

Depth EstimationRobot Manipulation

TediGAN: Text-Guided Diverse Face Image Generation and Manipulation

2020-12-06 · CVPR 2021 1 · Weihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan Wu

In this work, we propose TediGAN, a novel framework for multi-modal image generation and manipulation with textual descriptions. The proposed method consists of three components: StyleGAN inversion module, visual-linguis…

Face Sketch SynthesisImage GenerationText-to-Image Generation

LDEdit: Towards Generalized Text Guided Image Manipulation via Latent Diffusion Models

2022-10-05 · Paramanand Chandramouli, Kanchana Vaishnavi Gandikota

Research in vision-language models has seen rapid developments off-late, enabling natural language-based interfaces for image generation and manipulation. Many existing text guided manipulation techniques are restricted …

Image GenerationImage ManipulationStyle TransferText to Image Generation+1

DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation

2021-10-06 · CVPR 2022 1 · Gwanghyun Kim, Taesung Kwon, Jong Chul Ye

Recently, GAN inversion methods combined with Contrastive Language-Image Pretraining (CLIP) enables zero-shot image manipulation guided by text prompts. However, their applications to diverse real images are still diffic…

AttributeImage GenerationImage ManipulationTranslation