paper-with-me

Papers

TFCustom: Customized Image Generation with Time-Aware Frequency Feature Guidance

2025-01-01 · CVPR 2025 1 · Mushui Liu, Dong She, Jingxuan Pang, Qihan Huang, Jiacheng Ying, Wanggui He, Yuanlei Hou, Siming Fu

Subject-driven image personalization has seen notable advancements, especially with the advent of the ReferenceNet paradigm. ReferenceNet excels in integrating image reference features, making it highly applicable in creative and commercial settings. However, current implementations of ReferenceNet primarily operate as latent-level feature extractors, which limit their potential. This constraint hinders the provision of appropriate features to the denoising backbone across different timesteps, leading to suboptimal image consistency. In this paper, we revisit the extraction of reference features and propose TFCustom, a model framework designed to focus on reference image features at different temporal steps and frequency levels. Specifically, we firstly propose synchronized ReferenceNet to extract reference image features while simultaneously optimizing noise injection and denoising for the reference image. We also propose a time-aware frequency feature refinement module that leverages high- and low-frequency filters, combined with time embeddings, to adaptively select the degree of reference feature injection. Additionally, to enhance the similarity between reference objects and the generated image, we introduce a novel reward-based loss that encourages greater alignment between the reference and generated images. Experimental results demonstrate state-of-the-art performance in both multi-object and single-object reference generation, with significant improvements in texture and textual detail generation over existing methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingImage Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

RelationBooth: Towards Relation-Aware Customized Object Generation

2024-10-30 · Qingyu Shi, Lu Qi, Jianzong Wu, Jinbin Bai 외

Customized image generation is crucial for delivering personalized content based on user-provided image prompts, aligning large-scale text-to-image diffusion models with individual needs. However, existing models often o…

Image GenerationObjectRelation

CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers

2025-02-10 · D. She, Mushui Liu, Jingxuan Pang, Jin Wang 외

Customized generation has achieved significant progress in image synthesis, yet personalized video generation remains challenging due to temporal inconsistencies and quality degradation. In this paper, we introduce Custo…

Image GenerationVideo Generation

Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering

2024-03-14 · Zeyu Liu, Weicong Liang, Zhanhao Liang, Chong Luo 외

Visual text rendering poses a fundamental challenge for contemporary text-to-image generation models, with the core problem lying in text encoder deficiencies. To achieve accurate text rendering, we identify two crucial …

Image GenerationText to Image GenerationText-to-Image Generation

ThemeStation: Generating Theme-Aware 3D Assets from Few Exemplars

2024-03-22 · Zhenwei Wang, Tengfei Wang, Gerhard Hancke, Ziwei Liu 외

Real-world applications often require a large gallery of 3D assets that share a consistent theme. While remarkable advances have been made in general 3D content creation from text or image, synthesizing customized 3D ass…

3D GenerationDiversityUnity

Still-Moving: Customized Video Generation without Customized Video Data

2024-07-11 · Hila Chefer, Shiran Zada, Roni Paiss, Ariel Ephrat 외

Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation i…

Video Generation