paper-with-me

Papers

Towards Transformer-Based Aligned Generation with Self-Coherence Guidance

2025-03-22 · CVPR 2025 1 · Shulei Wang, Wang Lin, Hai Huang, Hanting Wang, Sihang Cai, WenKang Han, Tao Jin, Jingyuan Chen, Jiacheng Sun, Jieming Zhu, Zhou Zhao

We introduce a novel, training-free approach for enhancing alignment in Transformer-based Text-Guided Diffusion Models (TGDMs). Existing TGDMs often struggle to generate semantically aligned images, particularly when dealing with complex text prompts or multi-concept attribute binding challenges. Previous U-Net-based methods primarily optimized the latent space, but their direct application to Transformer-based architectures has shown limited effectiveness. Our method addresses these challenges by directly optimizing cross-attention maps during the generation process. Specifically, we introduce Self-Coherence Guidance, a method that dynamically refines attention maps using masks derived from previous denoising steps, ensuring precise alignment without additional training. To validate our approach, we constructed more challenging benchmarks for evaluating coarse-grained attribute binding, fine-grained attribute binding, and style binding. Experimental results demonstrate the superior performance of our method, significantly surpassing other state-of-the-art methods across all evaluated tasks. Our code is available at https://scg-diffusion.github.io/scg-diffusion.

📄 PDF Abstract BibTeX arXiv:2503.17675

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDenoising

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

2026-08-31 · Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar 외 arxiv

Clinical AI often optimizes predictive performance without engaging how clinicians decide where to look and what to write. We present Co-Annotator, which distills expert gaze and dictation into two guidance components: a…

Steering Video Diffusion Transformers with Massive Activations

2026-03-18 · Xianhang Cheng, Yujian Zheng, Zhenyu Xie, Tingting Liao 외 arxiv

Despite rapid progress in video diffusion transformers, how their internal model signals can be leveraged with minimal overhead to enhance video generation quality remains underexplored. In this work, we study the role o…

Video Generation

GeoDiff3D: Self-Supervised 3D Scene Generation with Geometry-Constrained 2D Diffusion Guidance

2026-01-27 · Haozhi Zhu, Miaomiao Zhao, Dingyao Liu, Runze Tian 외 arxiv

3D scene generation is a core technology for gaming, film/VFX, and VR/AR. Growing demand for rapid iteration, high-fidelity detail, and accessible content creation has further increased interest in this area. Existing me…

3D ReconstructionScene Generation3D Generation

Enhancing Object Coherence in Layout-to-Image Synthesis

2023-11-17 · Yibin Wang, Changhai Zhou, Honghui Xu

Layout-to-image synthesis is an emerging technique in conditional image generation. It aims to generate complex scenes, where users require fine control over the layout of the objects in a scene. However, it remains chal…

Conditional Image GenerationImage GenerationObjectTexture Synthesis

Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

2025-10-28 · Kang Zhang, Trung X. Pham, Suyeon Lee, Axi Niu 외 arxiv

We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier…

Audio Generation