paper-with-me

Papers

MIVE: New Design and Benchmark for Multi-Instance Video Editing

2024-12-17 · Samuel Teodoro, Agus Gunawan, Soo Ye Kim, Jihyong Oh, Munchurl Kim

Recent AI-based video editing has enabled users to edit videos through simple text prompts, significantly simplifying the editing process. However, recent zero-shot video editing techniques primarily focus on global or single-object edits, which can lead to unintended changes in other parts of the video. When multiple objects require localized edits, existing methods face challenges, such as unfaithful editing, editing leakage, and lack of suitable evaluation datasets and metrics. To overcome these limitations, we propose a zero-shot $\textbf{M}$ulti-$\textbf{I}$nstance $\textbf{V}$ideo $\textbf{E}$diting framework, called MIVE. MIVE is a general-purpose mask-based framework, not dedicated to specific objects (e.g., people). MIVE introduces two key modules: (i) Disentangled Multi-instance Sampling (DMS) to prevent editing leakage and (ii) Instance-centric Probability Redistribution (IPR) to ensure precise localization and faithful editing. Additionally, we present our new MIVE Dataset featuring diverse video scenarios and introduce the Cross-Instance Accuracy (CIA) Score to evaluate editing leakage in multi-instance video editing tasks. Our extensive qualitative, quantitative, and user study evaluations demonstrate that MIVE significantly outperforms recent state-of-the-art methods in terms of editing faithfulness, accuracy, and leakage prevention, setting a new benchmark for multi-instance video editing. The project page is available at https://kaist-viclab.github.io/mive-site/

📄 PDF Abstract BibTeX arXiv:2412.12877

Code (0)

등록된 구현이 없습니다.

Tasks

Video Editing

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MiVE: Multiscale Vision-language features for reference-guided video Editing

2026-05-14 · Tong Wang, Meng Zou, Chengjing Wu, Xiaochao Qu 외 arxiv

Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserving original motion and unedited content…

MIVE: A Minimalist Integer Vector Engine for Softmax LayerNorm and RMSNorm Acceleration

2026-06-16 · Kosmas Alexandridis, Giorgos Dimitrakopoulos arxiv

The rapid growth of Large Language Models (LLMs) has intensified the need for specialized hardware accelerators that can satisfy stringent inference latency and power constraints. Although matrix multiplications dominate…

A Multi-intersection Vehicular Cooperative Control based on End-Edge-Cloud Computing

2020-12-01 · Mingzhi Jiang, Tianhao Wu, Zhe Wang, Yi Gong 외

Cooperative Intelligent Transportation Systems (C-ITS) will change the modes of road safety and traffic management, especially at intersections without traffic lights, namely unsignalized intersections. Existing research…

Cloud ComputingManagement

Multiple Instance-Based Video Anomaly Detection using Deep Temporal Encoding-Decoding

2020-07-03 · Ammar Mansoor Kamoona, Amirali Khodadadian Gosta, Alireza Bab-Hadiashar, Reza Hoseinnezhad

In this paper, we propose a weakly supervised deep temporal encoding-decoding solution for anomaly detection in surveillance videos using multiple instance learning. The proposed approach uses both abnormal and normal vi…

Anomaly DetectionAnomaly Detection In Surveillance VideosMultiple Instance LearningVideo Anomaly Detection

MAIN: Multi-Attention Instance Network for Video Segmentation

2019-04-11 · Juan Leon Alcazar, Maria A. Bravo, Ali K. Thabet, Guillaume Jeanneret 외

Instance-level video segmentation requires a solid integration of spatial and temporal information. However, current methods rely mostly on domain-specific information (online learning) to produce accurate instance-level…

One-shot visual object segmentationSegmentationVideo SegmentationVideo Semantic Segmentation