paper-with-me

홈 › Papers

Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization

2026-06-01 · Liyuan Ma, Xueji Fang, Guo-Jun Qi arxiv

Image customization learns target subjects from reference concept images and generates conditioned images per text prompts, mainly modifying styles or backgrounds. Prevailing methods adopt fine-tuning to pack diverse concept attributes into a unified latent embedding, yet entangled attributes hinder elimination of irrelevant disturbances from style and background. To address this issue, we propose Equilibrated Diffusion, a frequency-driven approach that disentangles tangled concept features for balanced customization and consistent text-visual matching. Unlike conventional methods learning full concepts with shared embeddings and unified tuning, our work utilizes the inherent link between image frequency components and semantics: low frequencies represent subject content and high frequencies correspond to styles. We decompose concepts in frequency space and optimize each embedding independently. This separate optimization enables the denoiser to capture style detached from subject identity and generalize better to unseen stylistic prompts. Merging multi-frequency embeddings preserves the model's original spatial customization ability. We further deploy mask-guided diffusion to restrict irrelevant background changes and boost text alignment. Residual Reference Attention (RRA) is inserted into spatial attention to retain subject structure and identity consistency. Experiments prove Equilibrated Diffusion exceeds mainstream baselines on subject fidelity and text adherence, verifying our method's superiority.

📄 PDF Abstract BibTeX arXiv:2606.02129

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HyFAD: Hybrid Time-Frequency Diffusion with Frequency-Aware Embedding for Time Series Imputation

2026-06-03 · Hongfan Gao, Wangmeng Shen, Bin Yang, Jilin Hu arxiv

Diffusion models have demonstrated strong performance in time series modeling due to their ability to progressively capture complex data distributions through iterative denoising. However, existing approaches struggle wi…

Disentangled Textual Priors for Diffusion-based Image Super-Resolution

2026-03-08 · Lei Jiang, Xin Liu, Xinze Tong, Zhiliang Li 외 arxiv

Image Super-Resolution (SR) aims to reconstruct high-resolution images from degraded low-resolution inputs. While diffusion-based SR methods offer powerful generative capabilities, their performance heavily depends on ho…

Image Super-Resolution

ProSpect: Prompt Spectrum for Attribute-Aware Personalization of Diffusion Models

2023-05-25 · Yuxin Zhang, WeiMing Dong, Fan Tang, Nisha Huang 외

Personalizing generative models offers a way to guide image generation with user-provided references. Current personalization methods can invert an object or concept into the textual conditioning space and compose new na…

AttributeDisentanglementImage GenerationText-to-Image Generation

RDSplat: Robust Watermarking for 3D Gaussian Splatting Against 2D and 3D Diffusion Editing

2025-12-07 · Longjie Zhao, Ziming Hong, Zhenyang Ren, Runnan Chen 외 arxiv

3D Gaussian Splatting (3DGS) has become a leading representation for high-fidelity 3D assets, yet protecting these assets via digital watermarking remains an open challenge. Existing 3DGS watermarking methods are robust …

FCDM: A Physics-Guided Bidirectional Frequency Aware Convolution and Diffusion-Based Model for Sinogram Inpainting

2024-08-26 · Jiaze E, Srutarshi Banerjee, Tekin Bicer, Guannan Wang 외

Computed tomography (CT) is widely used in industrial and medical imaging, but sparse-view scanning reduces radiation exposure at the cost of incomplete sinograms and challenging reconstruction. Existing RGB-based inpain…

Computed Tomography (CT)CT ReconstructionImage ReconstructionScheduling+1