paper-with-me

홈 › Papers

Decoupled Guidance: Disentangling Subject and Context Pathways in Text-to-Image Personalization

2026-07-01 · Seongmin Kim, Kyucheol Shin, Heesun Jung, Jinseo Kim, Sungyong Baik arxiv

Text-to-image personalization aims to generate a user-provided subject in novel scenes described by text. However, most existing methods encode subject identity (fidelity) and context (editability) through the same conditioning pathway, forcing the two to compete for attention-map resources. We refer to this phenomenon as conditioning entanglement and show that it induces a fidelity-editability trade-off. We further provide causal evidence by replacing the target subject token with a generic subject token, which produces shifts in attention allocation and corresponding changes in context adherence. To this end, we propose Decoupled Guidance (DeGu), a plug-and-play framework that routes subject identity and scene context through two independent guidance streams. We further introduce a spatial mixing mechanism that dynamically fuses these streams, ensuring each operates within its semantically relevant region without interference. Furthermore, DeGu can be readily applied to existing personalization methods without modifying the underlying backbone models, consistently improving the overall personalization performance while enabling inference-time control over the fidelity-editability balance, across diverse methods and backbones, including flow-matching Diffusion Transformers (DiTs).

📄 PDF Abstract BibTeX arXiv:2607.00766

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RealCustom++: Representing Images as Real-Word for Real-Time Customization

2024-08-19 · Zhendong Mao, Mengqi Huang, Fei Ding, Mingcong Liu 외

Text-to-image customization, which takes given texts and images depicting given subjects as inputs, aims to synthesize new images that align with both text semantics and subject appearance. This task provides precise con…

Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context Transformer

2024-04-26 · Xinpeng Li, Teng Wang, Jian Zhao, Shuyi Mao 외

Emotion recognition aims to discern the emotional state of subjects within an image, relying on subject-centric and contextual visual cues. Current approaches typically follow a two-stage pipeline: first localize subject…

Emotion ClassificationEmotion Recognition

EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance

2024-09-12 · Zicheng Duan, Yuxuan Ding, Chenhui Gou, Ziqin Zhou 외

Zero-shot personalized image generation models aim to produce images that align with both a given text prompt and subject image, requiring the model to effectively incorporate both sources of guidance. However, existing …

DenoisingImage GenerationPersonalized Image GenerationSubject Transfer

Grounded AI for Code Review: Resource-Efficient Large-Model Serving in Enterprise Pipelines

2025-10-11 · Sayan Mandal, Hua Jiang arxiv

Automated code review adoption lags in compliance-heavy settings, where static analyzers produce high-volume, low-rationale outputs, and naive LLM use risks hallucination and incurring cost overhead. We present a product…

OmniGen2: Exploration to Advanced Multimodal Generation

2025-06-23 · Chenyuan Wu, Pengfei Zheng, Ruiran Yan, Shitao Xiao 외

In this work, we introduce OmniGen2, a versatile and open-source generative model designed to provide a unified solution for diverse generation tasks, including text-to-image, image editing, and in-context generation. Un…

Image Generationmultimodal generationText Generation