DreamMotion: Space-Time Self-Similar Score Distillation for Zero-Shot Video Editing
Text-driven diffusion-based video editing presents a unique challenge not encountered in image editing literature: establishing real-world motion. Unlike existing video editing approaches, here we focus on score distillation sampling to circumvent the standard reverse diffusion process and initiate optimization from videos that already exhibit natural motion. Our analysis reveals that while video score distillation can effectively introduce new content indicated by target text, it can also cause significant structure and motion deviation. To counteract this, we propose to match space-time self-similarities of the original video and the edited video during the score distillation. Thanks to the use of score distillation, our approach is model-agnostic, which can be applied for both cascaded and non-cascaded video diffusion frameworks. Through extensive comparisons with leading methods, our approach demonstrates its superiority in altering appearances while accurately preserving the original structure and motion.
Code (0)
등록된 구현이 없습니다.
Tasks
Video EditingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Tracing metaphors in time through self-distance in vector spaces
From a diachronic corpus of Italian, we build consecutive vector spaces in time and use them to compare a term's cosine similarity to itself in different time spans. We assume that a drop in similarity might be related t…
Contrastive Similarity Learning for Market Forecasting: The ContraSim Framework
We introduce the Contrastive Similarity Space Embedding Algorithm (ContraSim), a novel framework for uncovering the global semantic relationships between daily financial headlines and market movements. ContraSim operates…
Contrastive LearningRead Your Own Mind: Reasoning Helps Surface Self-Confidence Signals in LLMs
We study the source of uncertainty in DeepSeek R1-32B by analyzing its self-reported verbal confidence on question answering (QA) tasks. In the default answer-then-confidence setting, the model is regularly over-confiden…
Question AnsweringParameter Sharing Methods for Multilingual Self-Attentional Translation Models
In multilingual neural machine translation, it has been shown that sharing a single translation model between multiple languages can achieve competitive performance, sometimes even leading to performance gains over bilin…
Machine TranslationTranslationLearning Self-Similarity in Space and Time as Generalized Motion for Video Action Recognition
Spatio-temporal convolution often fails to learn motion dynamics in videos and thus an effective motion representation is required for video understanding in the wild. In this paper, we propose a rich and robust motion r…
Action RecognitionTemporal Action LocalizationVideo Understanding