paper-with-me

홈 › Papers

VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing

2025-02-24 · Xiangpeng Yang, Linchao Zhu, Hehe Fan, Yi Yang

Recent advancements in diffusion models have significantly improved video generation and editing capabilities. However, multi-grained video editing, which encompasses class-level, instance-level, and part-level modifications, remains a formidable challenge. The major difficulties in multi-grained editing include semantic misalignment of text-to-region control and feature coupling within the diffusion model. To address these difficulties, we present VideoGrain, a zero-shot approach that modulates space-time (cross- and self-) attention mechanisms to achieve fine-grained control over video content. We enhance text-to-region control by amplifying each local prompt's attention to its corresponding spatial-disentangled region while minimizing interactions with irrelevant areas in cross-attention. Additionally, we improve feature separation by increasing intra-region awareness and reducing inter-region interference in self-attention. Extensive experiments demonstrate our method achieves state-of-the-art performance in real-world scenarios. Our code, data, and demos are available at https://knightyxp.github.io/VideoGrain_project_page/

📄 PDF Abstract BibTeX arXiv:2502.17258

Code (0)

등록된 구현이 없습니다.

Tasks

Video EditingVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Learning Self-Modulating Attention in Continuous Time Space with Applications to Sequential Recommendation

2022-03-30 · Chao Chen, Haoyu Geng, Nianzu Yang, Junchi Yan 외

User interests are usually dynamic in the real world, which poses both theoretical and practical challenges for learning accurate preferences from rich behavior data. Among existing user behavior modeling solutions, atte…

Dynamic Link PredictionSequential Recommendation

Virtual Classification: Modulating Domain-Specific Knowledge for Multidomain Crowd Counting

2024-02-06 · Mingyue Guo, Binghui Chen, Zhaoyi Yan, YaoWei Wang 외

Multidomain crowd counting aims to learn a general model for multiple diverse datasets. However, deep networks prefer modeling distributions of the dominant domains instead of all domains, which is known as domain bias. …

Crowd Counting

Mamba Modulation: On the Length Generalization of Mamba

2025-09-23 · Peng Lu, Jerry Huang, Qiuhao Zeng, Xinyu Wang 외 arxiv

The quadratic complexity of the attention mechanism in Transformer models has motivated the development of alternative architectures with sub-quadratic scaling, such as state-space models. Among these, Mamba has emerged …

Beamforming with Oversampled Time-Modulated Arrays

2025-01-30 · Marcin Wachowiak, André Bourdoux, Sofie Pollin

The time-modulated array (TMA) is a simple array architecture in which each antenna is connected via a multi-throw switch. The switch acts as a modulator switching state faster than the symbol rate. The phase shifting an…

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models

2025-09-18 · Qidong Wang, Junjie Hu, Ming Jiang arxiv

Recent advances in causal interpretability have extended from language models to vision-language models (VLMs), seeking to reveal their internal mechanisms through input interventions. While textual interventions often t…