CAMEL: CAusal Motion Enhancement Tailored for Lifting Text-driven Video Editing
Text-driven video editing poses significant challenges in exhibiting flicker-free visual continuity while preserving the inherent motion patterns of original videos. Existing methods operate under a paradigm where motion and appearance are intricately intertwined. This coupling leads to the network either over-fitting appearance content -- failing to capture motion patterns -- or focusing on motion patterns at the expense of content generalization to diverse textual scenarios. Inspired by the pivotal role of wavelet transform in dissecting video sequences we propose CAusal Motion Enhancement tailored for Lifting text-driven video editing (CAMEL) a novel technique with two core designs. First we introduce motion prompts designed to summarize motion concepts from video templates through direct optimization. The optimized prompts are purposefully integrated into latent representations of diffusion models to enhance the motion fidelity of generated results. Second to enhance motion coherence and extend the generalization of appearance content to creative textual prompts we propose the causal motion-enhanced attention mechanism. This mechanism is implemented in tandem with a novel causal motion filter synergistically enhancing the motion coherence of disentangled high-frequency components and concurrently preserving the generalization of appearance content across various textual scenarios. Extensive experimental results show the superior performance of CAMEL.
Code (0)
등록된 구현이 없습니다.
Tasks
Video EditingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
CWNet: Causal Wavelet Network for Low-Light Image Enhancement
Traditional Low-Light Image Enhancement (LLIE) methods primarily focus on uniform brightness adjustment, often neglecting instance-level semantic information and the inherent characteristics of different features. To add…
Low-Light Image EnhancementMetric LearningA Structure-aware and Motion-adaptive Framework for 3D Human Pose Estimation with Mamba
Recent Mamba-based methods for the pose-lifting task tend to model joint dependencies by 2D-to-1D mapping with diverse scanning strategies. Though effective, they struggle to model intricate joint connections and uniform…
3D Human Pose EstimationPhysics-Based Causal Lifting Linearization of Nonlinear Control Systems Underpinned by the Koopman Operator
Methods for constructing causal linear models from nonlinear dynamical systems through lifting linearization underpinned by Koopman operator and physical system modeling theory are presented. Outputs of a nonlinear contr…
A Confidence-based Acquisition Model for Self-supervised Active Learning and Label Correction
Supervised neural approaches are hindered by their dependence on large, meticulously annotated datasets, a requirement that is particularly cumbersome for sequential tasks. The quality of annotations tends to deteriorate…
Active LearningClinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding
We present Clinical Camel, an open large language model (LLM) explicitly tailored for clinical research. Fine-tuned from LLaMA-2 using QLoRA, Clinical Camel achieves state-of-the-art performance across medical benchmarks…
GPULanguage ModelingLanguage ModellingLarge Language Model+1