paper-with-me

Papers

U-VAP: User-specified Visual Appearance Personalization via Decoupled Self Augmentation

2024-03-29 · CVPR 2024 1 · You Wu, Kean Liu, Xiaoyue Mi, Fan Tang, Juan Cao, Jintao Li

Concept personalization methods enable large text-to-image models to learn specific subjects (e.g., objects/poses/3D models) and synthesize renditions in new contexts. Given that the image references are highly biased towards visual attributes, state-of-the-art personalization models tend to overfit the whole subject and cannot disentangle visual characteristics in pixel space. In this study, we proposed a more challenging setting, namely fine-grained visual appearance personalization. Different from existing methods, we allow users to provide a sentence describing the desired attributes. A novel decoupled self-augmentation strategy is proposed to generate target-related and non-target samples to learn user-specified visual attributes. These augmented data allow for refining the model's understanding of the target attribute while mitigating the impact of unrelated attributes. At the inference stage, adjustments are conducted on semantic space through the learned target and non-target embeddings to further enhance the disentanglement of target attributes. Extensive experiments on various kinds of visual attributes with SOTA personalization methods show the ability of the proposed method to mimic target visual appearance in novel contexts, thus improving the controllability and flexibility of personalization.

📄 PDF Abstract BibTeX arXiv:2403.20231

Code (1)

ictmcg/u-vap 공식 구현 pytorch

Tasks

AttributeDisentanglementSentence

Similar Papers 제목 키워드 기반

Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization

2024-01-30 · Henglei Lv, Jiayu Xiao, Liang Li, Qingming Huang

Diffusion-based text-to-image personalization have achieved great success in generating subjects specified by users among various contexts. Even though, existing finetuning-based methods still suffer from model overfitti…

Diversity

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation

2025-07-08 · Zhenghao Zhang, Junchao Liao, Xiangyu Meng, Long Qin 외

Recent advances in diffusion transformer models for motion-guided video generation, such as Tora, have shown significant progress. In this paper, we present Tora2, an enhanced version of Tora, which introduces several de…

Video Generation

TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts

2024-01-26 · Jingyu Zhuang, Di Kang, Yan-Pei Cao, Guanbin Li 외

Text-driven 3D scene editing has gained significant attention owing to its convenience and user-friendliness. However, existing methods still lack accurate control of the specified appearance and location of the editing …

3D scene Editing

IDDM: Identity-Decoupled Personalized Diffusion Models with a Tunable Privacy-Utility Trade-off

2026-04-01 · Linyan Dai, Xinwei Zhang, Haoyang Li, Qingqing Ye 외 arxiv

Personalized text-to-image diffusion models (e.g., DreamBooth, LoRA) enable users to synthesize high-fidelity avatars from a few reference photos for social expression. However, once these generations are shared on socia…

Face Recognition

V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping

2025-12-13 · Hyunkoo Lee, Wooseok Jang, Jini Yang, Taehwan Kim 외 arxiv

Video personalization aims to generate videos that faithfully reflect a user-provided subject while following a text prompt. However, existing approaches often rely on heavy video-based finetuning or large-scale video da…