paper-with-me

홈 › Papers

InstructVEdit: A Holistic Approach for Instructional Video Editing

2025-03-22 · Chi Zhang, Chengjian Feng, Feng Yan, Qiming Zhang, Mingjin Zhang, Yujie Zhong, Jing Zhang, Lin Ma

Video editing according to instructions is a highly challenging task due to the difficulty in collecting large-scale, high-quality edited video pair data. This scarcity not only limits the availability of training data but also hinders the systematic exploration of model architectures and training strategies. While prior work has improved specific aspects of video editing (e.g., synthesizing a video dataset using image editing techniques or decomposed video editing training), a holistic framework addressing the above challenges remains underexplored. In this study, we introduce InstructVEdit, a full-cycle instructional video editing approach that: (1) establishes a reliable dataset curation workflow to initialize training, (2) incorporates two model architectural improvements to enhance edit quality while preserving temporal consistency, and (3) proposes an iterative refinement strategy leveraging real-world data to enhance generalization and minimize train-test discrepancies. Extensive experiments show that InstructVEdit achieves state-of-the-art performance in instruction-based video editing, demonstrating robust adaptability to diverse real-world scenarios. Project page: https://o937-blip.github.io/InstructVEdit.

📄 PDF Abstract BibTeX arXiv:2503.17641

Code (0)

등록된 구현이 없습니다.

Tasks

Video Editing

Similar Papers 제목 키워드 기반

VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation

2025-03-18 · Shoubin Yu, Difan Liu, Ziqiao Ma, Yicong Hong 외

Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a unified framework. In this paper, we int…

Reasoning SegmentationVideo Editing

Region-Constraint In-Context Generation for Instructional Video Editing

2025-12-19 · Zhongwei Zhang, Fuchen Long, Wei Li, Zhaofan Qiu 외 arxiv

The In-context generation paradigm recently has demonstrated strong power in instructional image editing with both data efficiency and synthesis quality. Nevertheless, shaping such in-context learning for instruction-bas…

Image Editing

RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives

2024-05-28 · Jaehong Yoon, Shoubin Yu, Mohit Bansal

Recent video generative models primarily rely on carefully written text prompts for specific tasks, like inpainting or style editing. They require labor-intensive textual descriptions for input videos, hindering their fl…

AttributeVideo Editing

Script-to-Slide Grounding: Grounding Script Sentences to Slide Objects for Automatic Instructional Video Generation

2026-03-14 · Rena Suzuki, Masato Kikuchi, Tadachika Ozono arxiv

While slide-based videos augmented with visual effects are widely utilized in education and research presentations, the video editing process -- particularly applying visual effects to ground spoken content to slide obje…

Video Generation

Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions

2024-03-11 · Lan Wang, Vishnu Boddeti, SerNam Lim

We introduce a novel text-to-pose video editing method, ReimaginedAct. While existing video editing tasks are limited to changes in attributes, backgrounds, and styles, our method aims to predict open-ended human action …

counterfactualVideo EditingVideo Understanding