paper-with-me

홈 › Papers

VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing

2024-06-18 · Jing Gu, Yuwei Fang, Ivan Skorokhodov, Peter Wonka, Xinya Du, Sergey Tulyakov, Xin Eric Wang

Video editing serves as a fundamental pillar of digital media, spanning applications in entertainment, education, and professional communication. However, previous methods often overlook the necessity of comprehensively understanding both global and local contexts, leading to inaccurate and inconsistent edits in the spatiotemporal dimension, especially for long videos. In this paper, we introduce VIA, a unified spatiotemporal Video Adaptation framework for global and local video editing, pushing the limits of consistently editing minute-long videos. First, to ensure local consistency within individual frames, we designed test-time editing adaptation to adapt a pre-trained image editing model for improving consistency between potential editing directions and the text instruction, and adapts masked latent variables for precise local control. Furthermore, to maintain global consistency over the video sequence, we introduce spatiotemporal adaptation that recursively gather consistent attention variables in key frames and strategically applies them across the whole sequence to realize the editing effects. Extensive experiments demonstrate that, compared to baseline methods, our VIA approach produces edits that are more faithful to the source videos, more coherent in the spatiotemporal context, and more precise in local control. More importantly, we show that VIA can achieve consistent long video editing in minutes, unlocking the potential for advanced video editing tasks over long video sequences.

📄 PDF Abstract BibTeX arXiv:2406.12831

Code (0)

등록된 구현이 없습니다.

Tasks

Video Editing

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention

2026-03-23 · Junhao Du, Jialong Xue, Anqi Li, Jincheng Dai 외 arxiv

Video large language models (Video-LLMs) face high computational costs due to large volumes of visual tokens. Existing token compression methods typically adopt a two-stage spatiotemporal compression strategy, relying on…

Semantic SimilarityQuestion Answering

UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

2022-01-12 · Kunchang Li, Yali Wang, Peng Gao, Guanglu Song 외

It is a challenging task to learn rich and multi-scale spatiotemporal semantics from high-dimensional videos, due to large local redundancy and complex global dependency between video frames. The recent advances in this …

Representation Learning

MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion

2024-10-10 · Onkar Susladkar, Jishu Sen Gupta, Chirag Sehgal, Sparsh Mittal 외

The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video p…

Denoisingparameter-efficient fine-tuningQuantizationText-to-Video Generation+3

StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental Learning

2025-05-20 · Huaijie Wang, De Cheng, Guozhang Li, Zhipeng Xu 외

Video Class-Incremental Learning (VCIL) seeks to develop models that continuously learn new action categories over time without forgetting previously acquired knowledge. Unlike traditional Class-Incremental Learning (CIL…

class-incremental learningClass Incremental LearningExemplar-FreeIncremental Learning+1

AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling

2026-05-15 · Ziyang Mai, Yuyao Zhang, Yu-Wing Tai arxiv

Recent diffusion-based video generators have achieved remarkable visual fidelity and prompt controllability, yet scaling them to ultra-high-resolution (UHR) long videos remains prohibitively expensive. The difficulty is …

Video Generation