paper-with-me

Papers

Event-Customized Image Generation

2024-10-03 · Zhen Wang, Yilei Jiang, Dong Zheng, Jun Xiao, Long Chen

Customized Image Generation, generating customized images with user-specified concepts, has raised significant attention due to its creativity and novelty. With impressive progress achieved in subject customization, some pioneer works further explored the customization of action and interaction beyond entity (i.e., human, animal, and object) appearance. However, these approaches only focus on basic actions and interactions between two entities, and their effects are limited by insufficient ''exactly same'' reference images. To extend customized image generation to more complex scenes for general real-world applications, we propose a new task: event-customized image generation. Given a single reference image, we define the ''event'' as all specific actions, poses, relations, or interactions between different entities in the scene. This task aims at accurately capturing the complex event and generating customized images with various target entities. To solve this task, we proposed a novel training-free event customization method: FreeEvent. Specifically, FreeEvent introduces two extra paths alongside the general diffusion denoising process: 1) Entity switching path: it applies cross-attention guidance and regulation for target entity generation. 2) Event transferring path: it injects the spatial feature and self-attention maps from the reference image to the target image for event generation. To further facilitate this new task, we collected two evaluation benchmarks: SWiG-Event and Real-Event. Extensive experiments and ablations have demonstrated the effectiveness of FreeEvent.

📄 PDF Abstract BibTeX arXiv:2410.02483

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingImage Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Non-confusing Generation of Customized Concepts in Diffusion Models

2024-05-11 · Wang Lin, Jingyuan Chen, Jiaxin Shi, Yichen Zhu 외

We tackle the common challenge of inter-concept visual confusion in compositional concept generation using text-guided diffusion models (TGDMs). It becomes even more pronounced in the generation of customized concepts, d…

RelationBooth: Towards Relation-Aware Customized Object Generation

2024-10-30 · Qingyu Shi, Lu Qi, Jianzong Wu, Jinbin Bai 외

Customized image generation is crucial for delivering personalized content based on user-provided image prompts, aligning large-scale text-to-image diffusion models with individual needs. However, existing models often o…

Image GenerationObjectRelation

Infusion: Preventing Customized Text-to-Image Diffusion from Overfitting

2024-04-22 · Weili Zeng, Yichao Yan, Qi Zhu, Zhuo Chen 외

Text-to-image (T2I) customization aims to create images that embody specific visual concepts delineated in textual descriptions. However, existing works still face a main challenge, concept overfitting. To tackle this ch…

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

2025-09-07 · Yi Yuan, Xubo Liu, Haohe Liu, Xiyuan Kang 외 arxiv

With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-…

Audio Generation

Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attacks

2025-03-11 · CVPR 2025 1 · Junying Wang, Hongyuan Zhang, Yuan Yuan

Recent Customized Portrait Generation (CPG) methods, taking a facial image and a textual prompt as inputs, have attracted substantial attention. Although these methods generate high-fidelity portraits, they fail to preve…

Face Recognition