paper-with-me

홈 › Papers

AttnRouter: Per-Category Attention Routing for Training-Free Image Editing on MMDiT

2026-05-02 · Guandong Li, Mengxia Ye arxiv

We study training-free image editing on Qwen-Image-Edit-2511, a 60-block multi-modal diffusion transformer (MMDiT) that concatenates noise and source-image tokens within a single attention stream. We make three contributions. (i) We introduce KVInject, a single-forward attention manipulation that alpha-blends source-half key/value projections into the noise-half within a localized layer/step band. KVInject is simpler than the classical two-pass MasaCtrl recipe and avoids the prompt-mismatch failure mode that disables MasaCtrl on MMDiT (composite score drops 31% versus baseline). (ii) We show that no single attention operation dominates across edit types, motivating AttnRouter, a per-category routing table that dispatches edits to the operation that best preserves source structure for that type. With ground-truth categories the router improves the CLIP-T+DINO-I composite by 6.4% over the editing baseline; an automatic CLIP zero-shot classifier closes 98% of this gap despite only 55% category accuracy. (iii) Through layer-, step-, and alpha-band ablations we localize the editing-effective attention sub-circuit: K/V injection in early denoising steps (S0-7) recovers nearly all of the gain of full-step injection, while injection in early (L0-15) or late (L45-60) layer bands fails to drive editing entirely; alpha in [0.3, 0.5] is a stable sweet spot. We also report negative results that highlight what does not transfer from the UNet folklore: simple K/V rescaling never beats baseline and aggressive variants collapse generation entirely (composite 0.084). We release code, pre-computed routing tables, and a 100-sample stratified subset of ImgEdit-Bench used in all ablations.

📄 PDF Abstract BibTeX arXiv:2605.01480

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference

2026-05-27 · Alan Ferrari arxiv

Standard transformer architectures apply a single attention mechanism uniformly across all tokens and sequence positions, irrespective of local context or computational budget. We propose Meta-Attention, a framework that…

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

2026-07-17 · Hao Liu, Chenghuan Huang, Ye Huang, Zhiying Wen 외 arxiv

Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p rou…

Temporal SequencesVideo Generation

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

2026-07-27 · Haopeng Li, Yitong Li, Junsong Chen, Tian Ye 외 hf

Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by compu…

Video Generation

HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads

2024-11-22 · Yu Xu, Fan Tang, Juan Cao, Yuxin Zhang 외

Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based…

Image Generationtext-guided-image-editing

FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time

2025-10-27 · Yaoli Liu, Yao-Xiang Ding, Kun Zhou arxiv

This paper proposes FreeFuse, a training-free framework for multi-subject text-to-image generation through automatic fusion of multiple subject LoRAs. In contrast to prior studies that focus on retraining LoRAs to allevi…

Text-to-Image Generation