paper-with-me

Papers

ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision

2025-12-24 · Weiqi Li, Zehao Zhang, Liang Lin, Guangrun Wang arxiv

Controllability is a fundamental requirement in video synthesis, where accurate alignment with conditioning signals is essential. Existing classifier-free guidance methods typically achieve conditioning indirectly by modeling the joint distribution of data and conditions, which often results in limited controllability over the specified conditions. Classifier-based guidance enforces conditions through an external classifier, but the model may exploit this mechanism to raise the classifier score without genuinely satisfying the intended condition, resulting in adversarial artifacts and limited effective controllability. In this paper, we propose Attention-Conditional Diffusion (ACD), a novel framework for direct conditional control in video diffusion models via attention supervision. By aligning the model's attention maps with external control signals, ACD achieves better controllability. To support this, we introduce a sparse 3D-aware object layout as an efficient conditioning signal, along with a dedicated Layout ControlNet and an automated annotation pipeline for scalable layout integration. Extensive experiments on benchmark video generation datasets demonstrate that ACD delivers superior alignment with conditioning inputs while preserving temporal coherence and visual fidelity, establishing an effective paradigm for conditional video synthesis.

📄 PDF Abstract BibTeX arXiv:2512.21268

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation

2023-10-11 · Bo Peng, Xinyuan Chen, Yaohui Wang, Chaochao Lu 외

Recent works have successfully extended large-scale text-to-image models to the video domain, producing promising results but at a high computational cost and requiring a large amount of video data. In this work, we intr…

Image GenerationText to Image GenerationText-to-Image GenerationText-to-Video Generation+1

Video-P2P: Video Editing with Cross-attention Control

2023-03-08 · CVPR 2024 1 · Shaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin 외

This paper presents Video-P2P, a novel framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there a…

Image GenerationVideo EditingVideo Generation

Video-P2P: Video Editing with Cross-attention Control

2023-03-08 · arXiv 2023 3 · Shaoteng Liu;Yuechen Zhang;Wenbo Li;Zhe Lin;Jiaya Jia

This paper presents Video-P2P, a novel framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there a…

Image GenerationVideo EditingVideo Generation

FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models

2024-06-24 · Haonan Qiu, Zhaoxi Chen, Zhouxia Wang, Yingqing He 외

Diffusion model has demonstrated remarkable capability in video generation, which further sparks interest in introducing trajectory control into the generation process. While existing works mainly focus on training-based…

Video Generation

DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning

2026-03-30 · Weimin Liu, Qingkun Li, Jiyuan Qiu, Wenjun Wang 외 arxiv

Drivers' visual attention provides critical cues for anticipating latent hazards and directly shapes decision-making and control maneuvers, where its absence can compromise traffic safety. To emulate drivers' perception …

Scene Understanding