paper-with-me

홈 › Papers

Cut-and-Paste: Subject-Driven Video Editing with Attention Control

2023-11-20 · Zhichao Zuo, Zhao Zhang, Yan Luo, Yang Zhao, Haijun Zhang, Yi Yang, Meng Wang

This paper presents a novel framework termed Cut-and-Paste for real-word semantic video editing under the guidance of text prompt and additional reference image. While the text-driven video editing has demonstrated remarkable ability to generate highly diverse videos following given text prompts, the fine-grained semantic edits are hard to control by plain textual prompt only in terms of object details and edited region, and cumbersome long text descriptions are usually needed for the task. We therefore investigate subject-driven video editing for more precise control of both edited regions and background preservation, and fine-grained semantic generation. We achieve this goal by introducing an reference image as supplementary input to the text-driven video editing, which avoids racking your brain to come up with a cumbersome text prompt describing the detailed appearance of the object. To limit the editing area, we refer to a method of cross attention control in image editing and successfully extend it to video editing by fusing the attention map of adjacent frames, which strikes a balance between maintaining video background and spatio-temporal consistency. Compared with current methods, the whole process of our method is like `cut" the source object to be edited and then `paste" the target object provided by reference image. We demonstrate that our method performs favorably over prior arts for video editing under the guidance of text prompt and extra reference image, as measured by both quantitative and subjective evaluations.

📄 PDF Abstract BibTeX arXiv:2311.11697

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectVideo Editing

Similar Papers 제목 키워드 기반

Paste, Inpaint and Harmonize via Denoising: Subject-Driven Image Editing with Pre-Trained Diffusion Model

2023-06-13 · Xin Zhang, Jiaxian Guo, Paul Yoo, Yutaka Matsuo 외

Text-to-image generative models have attracted rising attention for flexible image editing via user-specified descriptions. However, text descriptions alone are not enough to elaborate the details of subjects, often comp…

DenoisingImage GenerationScene Generation

DIVE: Taming DINO for Subject-Driven Video Editing

2024-12-04 · Yi Huang, Wei Xiong, He Zhang, Chaoqi Chen 외

Building on the success of diffusion models in image generation and editing, video editing has recently gained substantial attention. However, maintaining temporal consistency and motion alignment still remains challengi…

Image GenerationVideo Editing

ASTRA: Let Arbitrary Subjects Transform in Video Editing

2025-10-01 · Fei Shen, Weihao Xu, Rui Yan, Dong Zhang 외 arxiv

While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention dilution and mask boundary entanglement that cause attribute leakage and …

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

2026-04-26 · Haojie Zhang, Di Wu, Bingyan Liu, Linjie Zhong 외 arxiv

While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained by the absence of datasets that address…

Video GenerationVideo Alignment

VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment

2024-08-21 · Shangkun Sun, Xiaoyu Liang, Songlin Fan, Wenxu Gao 외

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective q…

Video AlignmentVideo EditingVideo Quality AssessmentVisual Question Answering (VQA)