paper-with-me

Papers

MTV-Inpaint: Multi-Task Long Video Inpainting

2025-03-14 · Shiyuan Yang, Zheng Gu, Liang Hou, Xin Tao, Pengfei Wan, Xiaodong Chen, Jing Liao

Video inpainting involves modifying local regions within a video, ensuring spatial and temporal consistency. Most existing methods focus primarily on scene completion (i.e., filling missing regions) and lack the capability to insert new objects into a scene in a controllable manner. Fortunately, recent advancements in text-to-video (T2V) diffusion models pave the way for text-guided video inpainting. However, directly adapting T2V models for inpainting remains limited in unifying completion and insertion tasks, lacks input controllability, and struggles with long videos, thereby restricting their applicability and flexibility. To address these challenges, we propose MTV-Inpaint, a unified multi-task video inpainting framework capable of handling both traditional scene completion and novel object insertion tasks. To unify these distinct tasks, we design a dual-branch spatial attention mechanism in the T2V diffusion U-Net, enabling seamless integration of scene completion and object insertion within a single framework. In addition to textual guidance, MTV-Inpaint supports multimodal control by integrating various image inpainting models through our proposed image-to-video (I2V) inpainting mode. Additionally, we propose a two-stage pipeline that combines keyframe inpainting with in-between frame propagation, enabling MTV-Inpaint to effectively handle long videos with hundreds of frames. Extensive experiments demonstrate that MTV-Inpaint achieves state-of-the-art performance in both scene completion and object insertion tasks. Furthermore, it demonstrates versatility in derived applications such as multi-modal inpainting, object editing, removal, image object brush, and the ability to handle long videos. Project page: https://mtv-inpaint.github.io/.

📄 PDF Abstract BibTeX arXiv:2503.11412

Code (0)

등록된 구현이 없습니다.

Tasks

Image InpaintingObjectVideo Inpainting

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
U-Net 설명 없음
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

CIRI: Curricular Inactivation for Residue-aware One-shot Video Inpainting

2023-01-01 · ICCV 2023 1 · Weiying Zheng, Cheng Xu, Xuemiao Xu, Wenxi Liu 외

Video inpainting aims at filling in missing regions of a video. However, when dealing with dynamic scenes with camera or object movements, annotating the inpainting target becomes laborious and impractical. In this p…

Object TrackingVideo Inpainting

Short-Term and Long-Term Context Aggregation Network for Video Inpainting

2020-09-12 · ECCV 2020 8 · Ang Li, Shanshan Zhao, Xingjun Ma, Mingming Gong 외

Video inpainting aims to restore missing regions of a video and has many applications such as video editing and object removal. However, existing methods either suffer from inaccurate short-term context aggregation or ra…

Video EditingVideo Inpainting

Towards Language-Driven Video Inpainting via Multimodal Large Language Models

2024-01-18 · CVPR 2024 1 · Jianzong Wu, Xiangtai Li, Chenyang Si, Shangchen Zhou 외

We introduce a new task -- language-driven video inpainting, which uses natural language instructions to guide the inpainting process. This approach overcomes the limitations of traditional video inpainting methods that …

Video Inpainting

Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising

2025-11-05 · Shuangquan Lyu, Steven Mao, Yue Ma arxiv

Generating long videos remains a fundamental challenge, and achieving high controllability in video inpainting and outpainting is particularly demanding. To address both of these challenges simultaneously and achieve con…

Video GenerationVideo Inpainting

DVI: Depth Guided Video Inpainting for Autonomous Driving

2020-07-17 · ECCV 2020 8 · Miao Liao, Feixiang Lu, Dingfu Zhou, Sibo Zhang 외

To get clear street-view and photo-realistic simulation in autonomous driving, we present an automatic video inpainting algorithm that can remove traffic agents from videos and synthesize missing regions with the guidanc…

Autonomous DrivingImage InpaintingPoint Cloud RegistrationVideo Inpainting