paper-with-me

Papers

OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing

2025-12-08 · Haoyang He, Jie Wang, Jiangning Zhang, Zhucun Xue, Xingyuan Bu, Qiangpeng Yang, Shilei Wen, Lei Xie arxiv

The quality and diversity of instruction-based image editing datasets are continuously increasing, yet large-scale, high-quality datasets for instruction-based video editing remain scarce. To address this gap, we introduce OpenVE-3M, an open-source, large-scale, and high-quality dataset for instruction-based video editing. It comprises two primary categories: spatially-aligned edits (Global Style, Background Change, Local Change, Local Remove, Local Add, and Subtitles Edit) and non-spatially-aligned edits (Camera Multi-Shot Edit and Creative Edit). All edit types are generated via a meticulously designed data pipeline with rigorous quality filtering. OpenVE-3M surpasses existing open-source datasets in terms of scale, diversity of edit types, instruction length, and overall quality. Furthermore, to address the lack of a unified benchmark in the field, we construct OpenVE-Bench, containing 431 video-edit pairs that cover a diverse range of editing tasks with three key metrics highly aligned with human judgment. We present OpenVE-Edit, a 5B model trained on our dataset that demonstrates remarkable efficiency and effectiveness by setting a new state-of-the-art on OpenVE-Bench, outperforming all prior open-source models including a 14B baseline. Project page is at https://lewandofskee.github.io/projects/OpenVE.

📄 PDF Abstract BibTeX arXiv:2512.07826

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance

2026-05-07 · Ziyun Zeng, Yiqi Lin, Guoqiang Liang, Mike Zheng Shou arxiv

In recent years, open-source efforts like Senorita-2M have propelled video editing toward natural language instruction. However, current publicly available datasets predominantly focus on local editing or style transfer,…

Style Transfer

OpenVeinNet: Robust Open-Set Finger Vein Verification with Dynamic Snake Convolution and Graph Learning

2026-08-26 · Sushrut Patwardhan, Raghavendra Ramachandra arxiv

Finger vein verification is a promising biometric modality for secure authentication because vascular patterns are internal, difficult to observe externally, and relatively resistant to presentation attacks. However, rel…

Graph Learning

LiveMoments: Reselected Key Photo Restoration in Live Photos via Reference-guided Diffusion

2026-04-14 · Clara Xue, Zizheng Yan, Zhenning Shi, Yuhang Yu 외 arxiv

Live Photo captures both a high-quality key photo and a short video clip to preserve the precious dynamics around the captured moment. While users may choose alternative frames as the key photo to capture better expressi…

Image Restoration

VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo Retouching

2026-04-30 · Yihong Guo, Youwei Lyu, Jiajun Tang, Yizhuo Zhou 외 arxiv

Reasoning photo retouching has gained significant traction, requiring models to analyze image defects, give reasoning processes, and execute precise retouching enhancements. However, existing approaches often rely on non…

Reinforcement LearningPhoto Retouching

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE

2026-05-04 · Yangming Shi, Shixiang Zhu, Tao Shen, Zhimiao Yu 외 arxiv

We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the model's generation capability, we equip t…

Reinforcement Learning