paper-with-me

Papers

Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions

2024-03-25 · CVPR 2025 1 · Stefan Andreas Baumann, Felix Krause, Michael Neumayr, Nick Stracke, Vincent Tao Hu, Björn Ommer

In recent years, advances in text-to-image (T2I) diffusion models have substantially elevated the quality of their generated images. However, achieving fine-grained control over attributes remains a challenge due to the limitations of natural language prompts (such as no continuous set of intermediate descriptions existing between `person'' and `old person''). Even though many methods were introduced that augment the model or generation process to enable such control, methods that do not require a fixed reference image are limited to either enabling global fine-grained attribute expression control or coarse attribute expression control localized to specific subjects, not both simultaneously. We show that there exist directions in the commonly used token-level CLIP text embeddings that enable fine-grained subject-specific control of high-level attributes in text-to-image models. Based on this observation, we introduce one efficient optimization-free and one robust optimization-based method to identify these directions for specific attributes from contrastive text prompts. We demonstrate that these directions can be used to augment the prompt text input with fine-grained control over attributes of specific subjects in a compositional manner (control over multiple attributes of a single subject) without having to adapt the diffusion model. Project page: https://compvis.github.io/attribute-control. Code is available at https://github.com/CompVis/attribute-control.

📄 PDF Abstract BibTeX arXiv:2403.17064

Code (1)

compvis/attribute-control 공식 구현 jax

Tasks

Attribute

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Ranking CGANs: Subjective Control over Semantic Image Attributes

2018-04-11 · Yassir Saquil, Kwang In Kim, Peter Hall

In this paper, we investigate the use of generative adversarial networks in the task of image generation according to subjective measures of semantic attributes. Unlike the standard (CGAN) that generates images from disc…

AttributeImage Generation

XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation

2025-06-26 · Bowen Chen, Mengyi Zhao, Haomiao Sun, Li Chen 외

Achieving fine-grained control over subject identity and semantic attributes (pose, style, lighting) in text-to-image generation, particularly for multiple subjects, often undermines the editability and coherence of Diff…

AttributeImage GenerationScene GenerationText to Image Generation+1

CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization

2024-09-09 · Nan Chen, Mengqi Huang, Zhuowei Chen, Yang Zheng 외

Subject-driven text-to-image (T2I) customization has drawn significant interest in academia and industry. This task enables pre-trained models to generate novel images based on unique subjects. Existing studies adopt a s…

Contrastive Learning

AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models

2025-08-04 · Die Chen, Zhongjie Duan, Zhiwen Li, Cen Chen 외 arxiv

Diffusion models have recently become the dominant paradigm for image generation, yet existing systems struggle to interpret and follow numeric instructions for adjusting semantic attributes. In real-world creative scena…

Reinforcement LearningContinuous ControlImage Generation

Constructing High Precision Knowledge Bases with Subjective and Factual Attributes

2019-05-28 · Ari Kobren, Pablo Barrio, Oksana Yakhnenko, Johann Hibschman 외

Knowledge bases (KBs) are the backbone of many ubiquitous applications and are thus required to exhibit high precision. However, for KBs that store subjective attributes of entities, e.g., whether a movie is "kid friendl…

AttributeVocal Bursts Intensity Prediction