Proposal-based Video Completion
Video inpainting is an important technique for a wide variety of applications from video content editing to video restoration. Early approaches follow image inpainting paradigms, but are challenged by complex camera motion and non-rigid deformations. To address these challenges flow-guided propagation techniques have been proposed. However, computation of flow is non-trivial for unobserved regions and propagation across a whole video sequence is computationally demanding. In contrast, in this paper, we propose a video inpainting algorithm based on proposals: we use 3D convolutions to obtain an initial inpainting estimate which is subsequently refined by fusing a generated set of proposals. Different from existing approaches for video inpainting, and inspired by well-explored mechanisms for object detection, we argue that proposals provide a rich source of information that permits to combine similarly looking patches that may be spatially and temporally far from the region to be inpainted. We validate the effectiveness of our method on the challenging YouTube VOS and DAVIS datasets using different settings and demonstrate results outperforming state-of-the-art on standard metrics.
Code (0)
등록된 구현이 없습니다.
Tasks
Image Inpaintingobject-detectionObject DetectionOne-shot visual object segmentationVideo InpaintingVideo RestorationSimilar Papers 제목 키워드 기반
Weakly-Supervised Video Moment Retrieval via Semantic Completion Network
Video moment retrieval is to search the moment that is most relevant to the given natural language query. Existing methods are mostly trained in a fully-supervised setting, which requires the full annotations of temporal…
Moment RetrievalRetrievalSemantic SimilaritySemantic Textual SimilarityIPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals
Semantic Scene Completion (SSC) has emerged as a pivotal approach for jointly learning scene geometry and semantics, enabling downstream applications such as navigation in mobile robotics. The recent generalization to Pa…
Scene UnderstandingWhat If We Do Not Have Multiple Videos of the Same Action? -- Video Action Localization Using Web Images
This paper tackles the problem of spatio-temporal action localization in a video without assuming the availability of multiple videos or any prior annotations. Action is localized by employing images downloaded from int…
Action LocalizationOptical Flow EstimationSpatio-Temporal Action LocalizationTemporal Action LocalizationMining Inter-Video Proposal Relations for Video Object Detection
Recent studies have shown that, context aggregating information from proposals in different frames can clearly enhance the performance of video object detection. However, these approaches mainly exploit the intra-proposa…
Objectobject-detectionObject DetectionRelation+3SCARP: 3D Shape Completion in ARbitrary Poses for Improved Grasping
Recovering full 3D shapes from partial observations is a challenging task that has been extensively addressed in the computer vision community. Many deep learning methods tackle this problem by training 3D shape generati…
3D Shape GenerationPose Estimationvalid