paper-with-me

Papers

StreamEdit: Training-Free Video Editing via Few-Step Streaming Video Generation

2026-05-20 · Guanlong Jiao, Chenyangguang Zhang, Jia Jun Cheng Xian, Zewei Zhang, Renjie Liao arxiv

Although existing video editing methods are generally feasible, they often require many costly iterations and still struggle to deliver high-quality yet satisfying editing results. We attribute this limitation to the prevalent data-to-data paradigm, which is less compatible with modern generative models than noise-to-data generation. To address this gap, we revisit video editing from a noise-to-data perspective and propose Streaming-Generation-based Video Editing (StreamEdit), which preserves few-step sampling while seamlessly injecting source-video conditions. Built on pre-trained streaming generation models, StreamEdit introduces dual-branch fast sampling with a self-attention bridge and cross-attention grounding/boosting to satisfy both sampling and conditioning requirements. We further propose source-oriented guidance to improve target-generation quality, and a visual prompting strategy to enhance editing flexibility and practicality. The method is effective, robust, and generalizable across different models. Extensive experiments on diverse video editing tasks show that StreamEdit consistently outperforms existing approaches, even in few-step settings with minimal time cost. Code and results are available at: https://dsl-lab.github.io/StreamEdit/.

📄 PDF Abstract BibTeX arXiv:2605.21466

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

2026-08-01 · Zhiqiang Lao arxiv

One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes such edits through low-energy smoothing along sampling time. Applie…

ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer

2026-03-16 · Ruonan Yu, Zhenxiong Tan, Zigeng Chen, Songhua Liu 외 arxiv

Diffusion Transformers (DiTs) have demonstrated remarkable scalability and quality in image and video generation, prompting growing interest in extending them to controllable generation and editing tasks. However, compar…

Video Generation

Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer

2025-08-12 · Zixin Yin, Xili Dai, Ling-Hao Chen, Deyu Zhou 외 arxiv

Text-guided color editing in images and videos is a fundamental yet unsolved problem, requiring fine-grained manipulation of color attributes, including albedo, light source color, and ambient lighting, while preserving …

Image Generation

V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes

2025-03-13 · YanMing Zhang, Jun-Kun Chen, Jipeng Lyu, Yu-Xiong Wang

This paper introduces V$^2$Edit, a novel training-free framework for instruction-guided video and 3D scene editing. Addressing the critical challenge of balancing original content preservation with editing task fulfillme…

3D scene EditingDenoisingVideo Editing

FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing

2024-09-30 · Lingling Cai, Kang Zhao, Hangjie Yuan, Yingya Zhang 외

Text-to-video diffusion models have made remarkable advancements. Driven by their ability to generate temporally coherent videos, research on zero-shot video editing using these fundamental models has expanded rapidly. T…

DenoisingVideo Editing