paper-with-me

홈 › Papers

Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model

2021-11-26 · CVPR 2022 1 · Zipeng Xu, Tianwei Lin, Hao Tang, Fu Li, Dongliang He, Nicu Sebe, Radu Timofte, Luc van Gool, Errui Ding

To achieve disentangled image manipulation, previous works depend heavily on manual annotation. Meanwhile, the available manipulations are limited to a pre-defined set the models were trained for. We propose a novel framework, i.e., Predict, Prevent, and Evaluate (PPE), for disentangled text-driven image manipulation that requires little manual annotation while being applicable to a wide variety of manipulations. Our method approaches the targets by deeply exploiting the power of the large-scale pre-trained vision-language model CLIP. Concretely, we firstly Predict the possibly entangled attributes for a given text command. Then, based on the predicted attributes, we introduce an entanglement loss to Prevent entanglements during training. Finally, we propose a new evaluation metric to Evaluate the disentangled image manipulation. We verify the effectiveness of our method on the challenging face editing task. Extensive experiments show that the proposed PPE framework achieves much better quantitative and qualitative results than the up-to-date StyleCLIP baseline.

📄 PDF Abstract BibTeX arXiv:2111.13333

Code (1)

zipengxuc/ppe 공식 구현 paddle

Tasks

Image ManipulationLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Training-Free Disentangled Text-Guided Image Editing via Sparse Latent Constraints

2025-12-25 · Mutiara Shabrina, Nova Kurnia Putri, Jefri Satria Ferdiansyah, Sabita Khansa Dewi 외 arxiv

Text-driven image manipulation often suffers from attribute entanglement, where modifying a target attribute (e.g., adding bangs) unintentionally alters other semantic properties such as identity or appearance. The Predi…

Image ManipulationImage GenerationImage Editing

Multi-input Architecture and Disentangled Representation Learning for Multi-dimensional Modeling of Music Similarity

2021-11-02 · Sebastian Ribecky, Jakob Abeßer, Hanna Lukashevich

In the context of music information retrieval, similarity-based approaches are useful for a variety of tasks that benefit from a query-by-example scenario. Music however, naturally decomposes into a set of semantically m…

DisentanglementInformation RetrievalMusic Information RetrievalRepresentation Learning+2

DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

2023-05-05 · Hong Chen, YiPeng Zhang, Simin Wu, Xin Wang 외

Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention. Existing methods mainly resort to finetuning a pretrai…

DenoisingDisentanglementImage GenerationText to Image Generation+1

Hierarchical Disentangled Representation Learning for Outdoor Illumination Estimation and Editing

2021-01-01 · ICCV 2021 10 · Piaopiao Yu, Jie Guo, Fan Huang, Cheng Zhou 외

Data-driven sky models have gained much attention in outdoor illumination prediction recently, showing superior performance against analytical models. However, naively compressing an outdoor panorama into a low-dimen…

Representation Learning

Designing User-Centric Behavioral Interventions to Prevent Dysglycemia with Novel Counterfactual Explanations

2023-10-02 · Asiful Arefeen, Hassan Ghasemzadeh

Monitoring unexpected health events and taking actionable measures to avert them beforehand is central to maintaining health and preventing disease. Therefore, a tool capable of predicting adverse health events and offer…

counterfactual