paper-with-me

홈 › Papers

Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation

2026-01-17 · Zijie Lou, Xiangwei Feng, Jiaxin Wang, Jiangtao Yao, Fei Che, Tianbao Liu, Chengjing Wu, Xiaochao Qu, Luoqi Liu, Ting Liu arxiv

Existing video object removal methods predominantly rely on diffusion models following a noise-to-data paradigm, where generation starts from uninformative Gaussian noise. This approach discards the rich structural and contextual priors present in the original input video. Consequently, such methods often lack sufficient guidance, leading to incomplete object erasure or the synthesis of implausible content that conflicts with the scene's physical logic. In this paper, we reformulate video object removal as a video-to-video translation task via a stochastic bridge model. Unlike noise-initialized methods, our framework establishes a direct stochastic path from the source video (with objects) to the target video (objects removed). This bridge formulation effectively leverages the input video as a strong structural prior, guiding the model to perform precise removal while ensuring that the filled regions are logically consistent with the surrounding environment. To address the trade-off where strong bridge priors hinder the removal of large objects, we propose a novel adaptive mask modulation strategy. This mechanism dynamically modulates input embeddings based on mask characteristics, balancing background fidelity with generative flexibility. Extensive experiments demonstrate that our approach significantly outperforms existing methods in both visual quality and temporal consistency. The project page is https://bridgeremoval.github.io/.

📄 PDF Abstract BibTeX arXiv:2601.12066

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal

2026-09-15 · Haonan Huang, Tianrui Qiu, Xianghao Zang, Yinan Du 외 arxiv

Despite its crucial role in video object removal (VOR), existing evaluation paradigms face two critical limitations: questionable references and a misalignment between tradi- tional metrics and human preference. To addre…

Video Generation

VORNet: Spatio-temporally Consistent Video Inpainting for Object Removal

2019-04-14 · Ya-Liang Chang, Zhe Yu Liu, Winston Hsu

Video object removal is a challenging task in video processing that often requires massive human efforts. Given the mask of the foreground object in each frame, the goal is to complete (inpaint) the object region and gen…

Image InpaintingObjectOne-shot visual object segmentationOptical Flow Estimation+3

EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing

2026-03-19 · Yang Fu, Yike Zheng, Ziyun Dai, Henghui Ding arxiv

Video object removal aims to eliminate dynamic target objects and their visual effects, such as deformation, shadows, and reflections, while restoring seamless backgrounds. Recent diffusion-based video inpainting and obj…

Video Inpainting

VOID: Video Object and Interaction Deletion

2026-04-02 · Saman Motamed, William Harvey, Benjamin Klein, Luc Van Gool 외 arxiv

Existing video object removal methods excel at inpainting content "behind" the object and correcting appearance-level artifacts such as shadows and reflections. However, when the removed object has more significant inter…

MiniMax-Remover: Taming Bad Noise Helps Video Object Removal

2025-05-30 · Bojia Zi, Weixuan Peng, Xianbiao Qi, Jianan Wang 외

Recent advances in video diffusion models have driven rapid progress in video editing techniques. However, video object removal, a critical subtask of video editing, remains challenging due to issues such as hallucinated…

Video EditingVideo Generation