paper-with-me

홈 › Papers

When to Lock Attention: Training-Free KV Control in Video Diffusion

2026-03-10 · Tianyi Zeng, Jincheng Gao, Tianyi Wang, Zijie Meng, Miao Zhang, Jun Yin, Haoyuan Sun, Junfeng Jiao, Christian Claudel, Junbo Tan, Xueqian Wang arxiv

Maintaining background consistency while enhancing foreground quality remains a core challenge in video editing. Injecting full-image information often leads to background artifacts, whereas rigid background locking severely constrains the model's capacity for foreground generation. To address this issue, we propose KV-Lock, a training-free framework tailored for DiT-based video diffusion models. Our core insight is that the hallucination metric (variance of denoising prediction) directly quantifies generation diversity, which is inherently linked to the classifier-free guidance (CFG) scale. Building upon this, KV-Lock leverages diffusion hallucination detection to dynamically schedule two key components: the fusion ratio between cached background key-values (KVs) and newly generated KVs, and the CFG scale. When hallucination risk is detected, KV-Lock strengthens background KV locking and simultaneously amplifies conditional guidance for foreground generation, thereby mitigating artifacts and improving generation fidelity. As a training-free, plug-and-play module, KV-Lock can be easily integrated into any pre-trained DiT-based models. Extensive experiments validate that our method outperforms existing approaches in improved foreground quality with high background fidelity across various video editing tasks.

📄 PDF Abstract BibTeX arXiv:2603.09657

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

2026-07-27 · Haopeng Li, Yitong Li, Junsong Chen, Tian Ye 외 hf

Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by compu…

Video Generation

Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

2026-06-29 · Xingyu Mou, Zijin Huang, Tianze Zhang, Yuxin Ma 외 arxiv

Block diffusion is the dominant approach for scaling discrete diffusion language models (dLLMs), as fixed-size blocks preserve parallel decoding while keeping quadratic attention costs tractable. Yet blockwise generation…

Zero-Painter: Training-Free Layout Control for Text-to-Image Synthesis

2024-06-06 · CVPR 2024 1 · Marianna Ohanyan, Hayk Manukyan, Zhangyang Wang, Shant Navasardyan 외

We present Zero-Painter, a novel training-free framework for layout-conditional text-to-image synthesis that facilitates the creation of detailed and controlled imagery from textual prompts. Our method utilizes object ma…

Conditional Text-to-Image SynthesisImage Generation

Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing

2025-04-14 · Taihang Hu, Linxuan Li, Kai Wang, Yaxing Wang 외

Text-to-image generation has seen groundbreaking advancements with diffusion models, enabling high-fidelity synthesis and precise image editing through cross-attention manipulation. Recently, autoregressive (AR) models h…

Image GenerationText to Image GenerationText-to-Image Generation

Exemplar-free Continual Learning of Vision Transformers via Gated Class-Attention and Cascaded Feature Drift Compensation

2022-11-22 · Marco Cotogni, Fei Yang, Claudio Cusano, Andrew D. Bagdanov 외

We propose a new method for exemplar-free class incremental training of ViTs. The main challenge of exemplar-free continual learning is maintaining plasticity of the learner without causing catastrophic forgetting of pre…

Continual LearningExemplar-Free