paper-with-me

Papers

VicEdit: Learning to Edit Videos from Visual In-Context Examples

2026-08-17 · Yuji Wang, Teng Hu, Yuheng Chen, Ran Yi, Han Feng, Weijian Cao, Chengjie Wang, Lizhuang Ma, Jiangning Zhang arxiv

Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair. To facilitate this paradigm, we curate VicEdit-400K, the first large-scale dataset for visual in-context video editing. We develop an automated pipeline to generate 400K high-quality samples across ten task types, ensuring superior visual fidelity and semantic consistency through multi-dimensional filtering. Leveraging this foundation, we introduce VicEdit, a unified framework to bridge visual and textual contexts. To adaptively extract editing semantics from heterogeneous references, we design Modality-Adaptive Semantic Distillation, which produces modality-specific semantic tokens from visual references. These tokens are then synergistically integrated with textual instructions through Dual-Context Injection, enabling the generation process to benefit from both visual and textual signals. Extensive evaluations on VicEditBench demonstrate that VicEdit achieves state-of-the-art performance across both basic instruction editing and visual in-context editing tasks, establishing visual in-context learning as a powerful and controllable paradigm for video editing.

📄 PDF Abstract BibTeX arXiv:2608.16745

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

In-Context Sync-LoRA for Portrait Video Editing

2025-12-02 · Sagi Polaczek, Or Patashnik, Ali Mahdavi-Amiri, Daniel Cohen-Or arxiv

Editing portrait videos is a challenging task that requires flexible yet precise control over a wide range of modifications, such as appearance changes, expression edits, or the addition of objects. The key difficulty li…

FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

2023-10-09 · Yuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen 외

Text-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts. A major challenge in this task is to ensure that all frames in the edited video are visually consistent. Most rec…

Optical Flow EstimationText-to-Video EditingVideo Editing

Text2LIVE: Text-Driven Layered Image and Video Editing

2022-04-05 · Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten 외

We present a method for zero-shot, text-driven appearance manipulation in natural images and videos. Given an input image or video and a target text prompt, our goal is to edit the appearance of existing objects (e.g., o…

Video Editing

VCoME: Verbal Video Composition with Multimodal Editing Effects

2024-07-05 · Weibo Gong, Xiaojie Jin, Xin Li, Dongliang He 외

Verbal videos, featuring voice-overs or text overlays, provide valuable content but present significant challenges in composition, especially when incorporating editing effects to enhance clarity and visual appeal. In th…

EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing

2024-12-13 · Umar Khalid, Chen Chen Umar Khalid, Hasan Iqbal, Azib Farooq 외

Editing complex visual content based on ambiguous instructions remains a challenging problem in vision-language modeling. While existing models can contextualize content, they often struggle to grasp the underlying inten…

Language ModelingLanguage ModellingMultimodal Reasoning