paper-with-me

홈 › Papers

Enhancing Text-to-Image Diffusion Transformer via Split-Text Conditioning

2025-05-25 · Yu Zhang, Jialei Zhou, Xinchen Li, Qi Zhang, Zhongwei Wan, Tianyu Wang, Duoqian Miao, Changwei Wang, Longbing Cao

Current text-to-image diffusion generation typically employs complete-text conditioning. Due to the intricate syntax, diffusion transformers (DiTs) inherently suffer from a comprehension defect of complete-text captions. One-fly complete-text input either overlooks critical semantic details or causes semantic confusion by simultaneously modeling diverse semantic primitive types. To mitigate this defect of DiTs, we propose a novel split-text conditioning framework named DiT-ST. This framework converts a complete-text caption into a split-text caption, a collection of simplified sentences, to explicitly express various semantic primitives and their interconnections. The split-text caption is then injected into different denoising stages of DiT-ST in a hierarchical and incremental manner. Specifically, DiT-ST leverages Large Language Models to parse captions, extracting diverse primitives and hierarchically sorting out and constructing these primitives into a split-text input. Moreover, we partition the diffusion denoising process according to its differential sensitivities to diverse semantic primitive types and determine the appropriate timesteps to incrementally inject tokens of diverse semantic primitive types into input tokens via cross-attention. In this way, DiT-ST enhances the representation learning of specific semantic primitive types across different stages. Extensive experiments validate the effectiveness of our proposed DiT-ST in mitigating the complete-text comprehension defect.

📄 PDF Abstract BibTeX arXiv:2505.19261

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingReading ComprehensionRepresentation Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Enhancing Privacy in ControlNet and Stable Diffusion via Split Learning

2024-09-13 · Dixi Yao

With the emerging trend of large generative models, ControlNet is introduced to enable users to fine-tune pre-trained models with their own data for various use cases. A natural question arises: how can we train ControlN…

Federated LearningImage GenerationPrivacy Preserving

TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers

2026-01-05 · Binglei Li, Mengping Yang, Zhiyu Tan, Junping Zhang 외 arxiv

Recent breakthroughs of transformer-based diffusion models, particularly with Multimodal Diffusion Transformers (MMDiT) driven models like FLUX and Qwen Image, have facilitated thrilling experiences in visual generation.…

Text-to-Image GenerationImage Editing

Reducing Texture Bias of Deep Neural Networks via Edge Enhancing Diffusion

2024-02-14 · Edgar Heinert, Matthias Rottmann, Kira Maag, Karsten Kahl

Convolutional neural networks (CNNs) for image processing tend to focus on localized texture patterns, commonly referred to as texture bias. While most of the previous works in the literature focus on the task of image c…

Adversarial RobustnessDomain Generalizationimage-classificationImage Classification+2

EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution

2025-05-08 · Haizhen Xie, Kunpeng Du, Qiangyu Yan, Sen Lu 외

Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. While T2I models have traditionally relied on U-Net architectures, recent a…

Blind Super-ResolutionImage RestorationIn-Context LearningSuper-Resolution

X-CAUNET: Cross-Color Channel Attention with Underwater Image-Enhancing Transformer

2024-03-18 · ICASSP 2024 3 · Alik Pramanick, Sandipan Sarma, Arijit Sur

Underwater image enhancement is essential to mitigate the environment-centric noise in images, such as haziness, color degradation, etc. With most existing works focused on processing an RGB image as a whole, the explici…

Image EnhancementSSIM