paper-with-me

홈 › Papers

EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing

2024-03-24 · Xiangpeng Yang, Linchao Zhu, Hehe Fan, Yi Yang

Current diffusion-based video editing primarily focuses on local editing (\textit{e.g.,} object/background editing) or global style editing by utilizing various dense correspondences. However, these methods often fail to accurately edit the foreground and background simultaneously while preserving the original layout. We find that the crux of the issue stems from the imprecise distribution of attention weights across designated regions, including inaccurate text-to-attribute control and attention leakage. To tackle this issue, we introduce EVA, a \textbf{zero-shot} and \textbf{multi-attribute} video editing framework tailored for human-centric videos with complex motions. We incorporate a Spatial-Temporal Layout-Guided Attention mechanism that leverages the intrinsic positive and negative correspondences of cross-frame diffusion features. To avoid attention leakage, we utilize these correspondences to boost the attention scores of tokens within the same attribute across all video frames while limiting interactions between tokens of different attributes in the self-attention layer. For precise text-to-attribute manipulation, we use discrete text embeddings focused on specific layout areas within the cross-attention layer. Benefiting from the precise attention weight distribution, EVA can be easily generalized to multi-object editing scenarios and achieves accurate identity mapping. Extensive experiments demonstrate EVA achieves state-of-the-art results in real-world scenarios. Full results are provided at https://knightyxp.github.io/EVA/

📄 PDF Abstract BibTeX arXiv:2403.16111

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeVideo Editing

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MAC: A Benchmark for Multiple Attributes Compositional Zero-Shot Learning

2024-06-18 · Shuo Xu, Sai Wang, Xinyue Hu, Yutian Lin 외

Compositional Zero-Shot Learning (CZSL) aims to learn semantic primitives (attributes and objects) from seen compositions and recognize unseen attribute-object compositions. Existing CZSL datasets focus on single attribu…

AttributeCompositional Zero-Shot LearningZero-Shot Learning

Zero-Shot Activity Recognition with Verb Attribute Induction

2017-07-29 · EMNLP 2017 9 · Rowan Zellers, Yejin Choi

In this paper, we investigate large-scale zero-shot activity recognition by modeling the visual and linguistic attributes of action verbs. For example, the verb "salute" has several properties, such as being a light move…

Activity RecognitionAttribute

Attribute Attention for Semantic Disambiguation in Zero-Shot Learning

2019-10-01 · ICCV 2019 10 · Yang Liu, Jishun Guo, Deng Cai, Xiaofei He

Zero-shot learning (ZSL) aims to accurately recognize unseen objects by learning mapping matrices that bridge the gap between visual information and semantic attributes. Previous works implicitly treat attributes equally…

AttributeZero-Shot Learning

Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class

2024-04-25 · Mazda Moayeri, Michael Rabbat, Mark Ibrahim, Diane Bouchacourt

Vision-language models enable open-world classification of objects without the need for any retraining. While this zero-shot paradigm marks a significant advance, even today's best models exhibit skewed performance when …

Diversityzero-shot-classificationZero-Shot Learning

Make an Omelette with Breaking Eggs: Zero-Shot Learning for Novel Attribute Synthesis

2021-11-28 · Yu-Hsuan Li, Tzu-Yin Chao, Ching-Chun Huang, Pin-Yu Chen 외

Most of the existing algorithms for zero-shot classification problems typically rely on the attribute-based semantic relations among categories to realize the classification of novel categories without observing any of t…

AttributeClassificationzero-shot-classificationZero-Shot Learning