paper-with-me

Papers

Point2Insert: Video Object Insertion via Sparse Point Guidance

2026-02-04 · Yu Zhou, Xiaoyan Yang, Bojia Zi, Lihan Zhang, Ruijie Sun, Weishi Zheng, Haibin Huang, Chi Zhang, Xuelong Li arxiv

This paper introduces Point2Insert, a sparse-point-based framework for flexible and user-friendly object insertion in videos, motivated by the growing popularity of accurate, low-effort object placement. Existing approaches face two major challenges: mask-based insertion methods require labor-intensive mask annotations, while instruction-based methods struggle to place objects at precise locations. Point2Insert addresses these issues by requiring only a small number of sparse points instead of dense masks, eliminating the need for tedious mask drawing. Specifically, it supports both positive and negative points to indicate regions that are suitable or unsuitable for insertion, enabling fine-grained spatial control over object locations. The training of Point2Insert consists of two stages. In Stage 1, we train an insertion model that generates objects in given regions conditioned on either sparse-point prompts or a binary mask. In Stage 2, we further train the model on paired videos synthesized by an object removal model, adapting it to video insertion. Moreover, motivated by the higher insertion success rate of mask-guided editing, we leverage a mask-guided insertion model as a teacher to distill reliable insertion behavior into the point-guided model. Extensive experiments demonstrate that Point2Insert consistently outperforms strong baselines and even surpasses models with $\times$10 more parameters.

📄 PDF Abstract BibTeX arXiv:2602.04167

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Controllable Video Object Insertion via Multiview Priors

2026-04-16 · Xia Qi, Peishan Cong, Yichen Yao, Ziyi Wang 외 arxiv

Video object insertion is a critical task for dynamically inserting new objects into existing environments. Previous video generation methods focus primarily on synthesizing entire scenes while struggling with ensuring c…

Video Generation

DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image

2025-03-13 · Qi Zhao, Zhan Ma, Pan Zhou

Recent developments in generative diffusion models have turned many dreams into realities. For video object insertion, existing methods typically require additional information, such as a reference video or a 3D asset of…

Object

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

2025-01-02 · Yuanpeng Tu, Hao Luo, Xi Chen, Sihui Ji 외

Despite significant advancements in video generation, inserting a given object into videos remains a challenging task. The difficulty lies in preserving the appearance details of the reference object and accurately model…

Talking Head GenerationVideo GenerationVirtual Try-on

SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion

2026-05-22 · Xinyu Chen, Yuyi Qian, Jiang Lin, Shenyi Wang 외 arxiv

Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hindered by a reliance on explicit motion e…

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework

2026-05-22 · Xiao Cao, Yansong Qu, Xiangzhen, Chang 외 arxiv

Mask-free video object insertion has emerged as a challenging task, requiring harmonious integration of reference objects into source videos. However, existing methods struggle when references exhibit severe stylistic do…

Video GenerationStyle Transfer