paper-with-me

홈 › Papers

Text-Driven Reasoning Video Editing via Reinforcement Learning on Digital Twin Representations

2025-11-18 · Yiqing Shen, Chenjia Li, Mathias Unberath arxiv

Text-driven video editing enables users to modify video content only using text queries. While existing methods can modify video content if explicit descriptions of editing targets with precise spatial locations and temporal boundaries are provided, these requirements become impractical when users attempt to conceptualize edits through implicit queries referencing semantic properties or object relationships. We introduce reasoning video editing, a task where video editing models must interpret implicit queries through multi-hop reasoning to infer editing targets before executing modifications, and a first model attempting to solve this complex task, RIVER (Reasoning-based Implicit Video Editor). RIVER decouples reasoning from generation through digital twin representations of video content that preserve spatial relationships, temporal trajectories, and semantic attributes. A large language model then processes this representation jointly with the implicit query, performing multi-hop reasoning to determine modifications, then outputs structured instructions that guide a diffusion-based editor to execute pixel-level changes. RIVER training uses reinforcement learning with rewards that evaluate reasoning accuracy and generation quality. Finally, we introduce RVEBenchmark, a benchmark of 100 videos with 519 implicit queries spanning three levels and categories of reasoning complexity specifically for reasoning video editing. RIVER demonstrates best performance on the proposed RVEBenchmark and also achieves state-of-the-art performance on two additional video editing benchmarks (VegGIE and FiVE), where it surpasses six baseline methods.

📄 PDF Abstract BibTeX arXiv:2511.14100

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

A Reinforcement Learning-Based Automatic Video Editing Method Using Pre-trained Vision-Language Model

2024-11-07 · Panwen Hu, Nan Xiao, Feifei Li, Yongquan Chen 외

In this era of videos, automatic video editing techniques attract more and more attention from industry and academia since they can reduce workloads and lower the requirements for human editors. Existing automatic editin…

Language ModelingLanguage ModellingReinforcement Learning (RL)Video Editing

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs

2025-05-26 · Juntong Wang, Jiarui Wang, Huiyu Duan, Guangtao Zhai 외

Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. …

BenchmarkingLarge Language ModelVideo EditingVideo Quality Assessment+1

VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment

2024-08-21 · Shangkun Sun, Xiaoyu Liang, Songlin Fan, Wenxu Gao 외

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective q…

Video AlignmentVideo EditingVideo Quality AssessmentVisual Question Answering (VQA)

Cut-and-Paste: Subject-Driven Video Editing with Attention Control

2023-11-20 · Zhichao Zuo, Zhao Zhang, Yan Luo, Yang Zhao 외

This paper presents a novel framework termed Cut-and-Paste for real-word semantic video editing under the guidance of text prompt and additional reference image. While the text-driven video editing has demonstrated remar…

ObjectVideo Editing

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

2026-07-28 · Katsuya Ogata, Zongshang Pang, Mayu Otani, Yuta Nakashima arxiv

Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes to convey. Benchmarks for a closely related task, video summarizati…

Video Summarization