paper-with-me

Papers

Cross-Image Attention for Zero-Shot Appearance Transfer

2023-11-06 · Yuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor, Daniel Cohen-Or

Recent advancements in text-to-image generative models have demonstrated a remarkable ability to capture a deep semantic understanding of images. In this work, we leverage this semantic knowledge to transfer the visual appearance between objects that share similar semantics but may differ significantly in shape. To achieve this, we build upon the self-attention layers of these generative models and introduce a cross-image attention mechanism that implicitly establishes semantic correspondences across images. Specifically, given a pair of images -- one depicting the target structure and the other specifying the desired appearance -- our cross-image attention combines the queries corresponding to the structure image with the keys and values of the appearance image. This operation, when applied during the denoising process, leverages the established semantic correspondences to generate an image combining the desired structure and appearance. In addition, to improve the output image quality, we harness three mechanisms that either manipulate the noisy latent codes or the model's internal representations throughout the denoising process. Importantly, our approach is zero-shot, requiring no optimization or training. Experiments show that our method is effective across a wide range of object categories and is robust to variations in shape, size, and viewpoint between the two input images.

📄 PDF Abstract BibTeX arXiv:2311.03335

Code (0)

등록된 구현이 없습니다.

Tasks

Appearance TransferDenoising

Similar Papers 제목 키워드 기반

Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment

2025-08-27 · Namu Kim, Wonbin Kweon, Minsoo Kim, Hwanjo Yu arxiv

We observe that zero-shot appearance transfer with large-scale image generation models faces a significant challenge: Attention Leakage. This challenge arises when the semantic mapping between two images is captured by t…

Image Generation

AnimateZero: Video Diffusion Models are Zero-Shot Image Animators

2023-12-06 · Jiwen Yu, Xiaodong Cun, Chenyang Qi, Yong Zhang 외

Large-scale text-to-video (T2V) diffusion models have great progress in recent years in terms of visual quality, motion and temporal consistency. However, the generation process is still a black box, where all attributes…

Image AnimationVideo Generation

Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing

2025-05-29 · Tongtong Su, Chengyu Wang, Jun Huang, Dongming Lu

Appearance editing according to user needs is a pivotal task in video editing. Existing text-guided methods often lead to ambiguities regarding user intentions and restrict fine-grained control over editing specific aspe…

Optical Flow EstimationVideo EditingVideo Restoration

Mask-guided cross-image attention for zero-shot in-silico histopathologic image generation with a diffusion model

2024-07-16 · Dominik Winter, Nicolas Triltsch, Marco Rosati, Anatoliy Shumilov 외

Creating in-silico data with generative AI promises a cost-effective alternative to staining, imaging, and annotating whole slide images in computational pathology. Diffusion models are the state-of-the-art solution for …

Appearance TransferImage Generationwhole slide images

DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis

2023-12-20 · CVPR 2024 1 · Yuming Gu, You Xie, Hongyi Xu, Guoxian Song 외

We present DiffPortrait3D, a conditional diffusion model that is capable of synthesizing 3D-consistent photo-realistic novel views from as few as a single in-the-wild portrait. Specifically, given a single RGB input, we …

Denoising