paper-with-me

Papers

OmniPrism: Learning Disentangled Visual Concept for Image Generation

2024-12-16 · Yangyang Li, Daqing Liu, Wu Liu, Allen He, Xinchen Liu, Yongdong Zhang, Guoqing Jin

Creative visual concept generation often draws inspiration from specific concepts in a reference image to produce relevant outcomes. However, existing methods are typically constrained to single-aspect concept generation or are easily disrupted by irrelevant concepts in multi-aspect concept scenarios, leading to concept confusion and hindering creative generation. To address this, we propose OmniPrism, a visual concept disentangling approach for creative image generation. Our method learns disentangled concept representations guided by natural language and trains a diffusion model to incorporate these concepts. We utilize the rich semantic space of a multimodal extractor to achieve concept disentanglement from given images and concept guidance. To disentangle concepts with different semantics, we construct a paired concept disentangled dataset (PCD-200K), where each pair shares the same concept such as content, style, and composition. We learn disentangled concept representations through our contrastive orthogonal disentangled (COD) training pipeline, which are then injected into additional diffusion cross-attention layers for generation. A set of block embeddings is designed to adapt each block's concept domain in the diffusion models. Extensive experiments demonstrate that our method can generate high-quality, concept-disentangled results with high fidelity to text prompts and desired concepts.

📄 PDF Abstract BibTeX arXiv:2412.12242

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementImage Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Visual Concepts Tokenization

2022-05-20 · Tao Yang, Yuwang Wang, Yan Lu, Nanning Zheng

Obtaining the human-like perception ability of abstracting visual concepts from concrete pixels has always been a fundamental and important target in machine learning research fields such as disentangled representation l…

Representation Learning

DiMBERT: Learning Vision-Language Grounded Representations with Disentangled Multimodal-Attention

2022-10-28 · Fenglin Liu, Xian Wu, Shen Ge, Xuancheng Ren 외

Vision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (a.k.a. V-L representations) is of paramo…

Image CaptioningLanguage ModelingLanguage ModellingSentence+2

UniVerse: A Unified Modulation Framework for Segmentation-Free,Disentangled Multi-Concept Personalization

2026-05-29 · Quynh Phung, Sandesh Ghimire, Minsi Hu, Chung-Chi Tsai 외 arxiv

Personalized visual understanding has advanced significantly, yet existing approaches struggle to localize and extract specific concepts when input images contain multiple objects. Many prior methods rely heavily on segm…

Learning Disentangled Prompts for Compositional Image Synthesis

2023-06-01 · Kihyuk Sohn, Albert Shaw, Yuan Hao, Han Zhang 외

We study domain-adaptive image synthesis, the problem of teaching pretrained image generative models a new style or concept from as few as one image to synthesize novel images, to better understand the compositional imag…

Domain AdaptationImage GenerationVisual Prompt Tuning

Attention Calibration for Disentangled Text-to-Image Personalization

2024-03-27 · CVPR 2024 1 · Yanbing Zhang, Mengping Yang, Qin Zhou, Zhe Wang

Recent thrilling progress in large-scale text-to-image (T2I) models has unlocked unprecedented synthesis quality of AI-generated content (AIGC) including image generation, 3D and video composition. Further, personalized …

Image GenerationNovel Concepts