paper-with-me

홈 › Papers

ExpressEdit: Video Editing with Natural Language and Sketching

2024-03-26 · Bekzat Tilekbay, Saelyne Yang, Michal Lewkowicz, Alex Suryapranata, Juho Kim

Informational videos serve as a crucial source for explaining conceptual and procedural knowledge to novices and experts alike. When producing informational videos, editors edit videos by overlaying text/images or trimming footage to enhance the video quality and make it more engaging. However, video editing can be difficult and time-consuming, especially for novice video editors who often struggle with expressing and implementing their editing ideas. To address this challenge, we first explored how multimodality$-$natural language (NL) and sketching, which are natural modalities humans use for expression$-$can be utilized to support video editors in expressing video editing ideas. We gathered 176 multimodal expressions of editing commands from 10 video editors, which revealed the patterns of use of NL and sketching in describing edit intents. Based on the findings, we present ExpressEdit, a system that enables editing videos via NL text and sketching on the video frame. Powered by LLM and vision models, the system interprets (1) temporal, (2) spatial, and (3) operational references in an NL command and spatial references from sketching. The system implements the interpreted edits, which then the user can iterate on. An observational study (N=10) showed that ExpressEdit enhanced the ability of novice video editors to express and implement their edit ideas. The system allowed participants to perform edits more efficiently and generate more ideas by generating edits based on user's multimodal edit commands and supporting iterations on the editing commands. This work offers insights into the design of future multimodal interfaces and AI-based pipelines for video editing.

📄 PDF Abstract BibTeX arXiv:2403.17693

Code (0)

등록된 구현이 없습니다.

Tasks

Video Editing

Similar Papers 제목 키워드 기반

ExpressEdit: Fast Editing of Stylized Facial Expressions with Diffusion Models in Photoshop

2026-04-03 · Kenan Tang, Jiasheng Guo, Jeffrey Lin, Yao Qin arxiv

Facial expressions of characters are a vital component of visual storytelling. While current AI image editing models hold promise for assisting artists in the task of stylized expression editing, these models introduce g…

Visual StorytellingImage Editing

Sketch Video Synthesis

2023-11-26 · Yudian Zheng, Xiaodong Cun, Menghan Xia, Chi-Man Pun

Understanding semantic intricacies and high-level concepts is essential in image sketch generation, and this challenge becomes even more formidable when applied to the domain of videos. To address this, we propose a nove…

Video Editing

SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation

2025-05-25 · Shenggan Cheng, Yuanxin Wei, Lansong Diao, Yong liu 외

Leveraging the diffusion transformer (DiT) architecture, models like Sora, CogVideoX and Wan have achieved remarkable progress in text-to-video, image-to-video, and video editing tasks. Despite these advances, diffusion-…

Video EditingVideo Generation

Sketch3DVE: Sketch-based 3D-Aware Scene Video Editing

2025-08-19 · Feng-Lin Liu, Shi-Yang Li, Yan-Pei Cao, Hongbo Fu 외 arxiv

Recent video editing methods achieve attractive results in style transfer or appearance modification. However, editing the structural content of 3D scenes in videos remains challenging, particularly when dealing with sig…

Style TransferImage Editing

EditDuet: A Multi-Agent System for Video Non-Linear Editing

2025-09-13 · Marcelo Sandoval-Castaneda, Bryan Russell, Josef Sivic, Gregory Shakhnarovich 외 arxiv

Automated tools for video editing and assembly have applications ranging from filmmaking and advertisement to content creation for social media. Previous video editing work has mainly focused on either retrieval or user …

Decision Making