Towards Unified Keyframe Propagation Models
Many video editing tasks such as rotoscoping or object removal require the propagation of context across frames. While transformers and other attention-based approaches that aggregate features globally have demonstrated great success at propagating object masks from keyframes to the whole video, they struggle to propagate high-frequency details such as textures faithfully. We hypothesize that this is due to an inherent bias of global attention towards low-frequency features. To overcome this limitation, we present a two-stream approach, where high-frequency features interact locally and low-frequency features interact globally. The global interaction stream remains robust in difficult situations such as large camera motions, where explicit alignment fails. The local interaction stream propagates high-frequency details through deformable feature aggregation and, informed by the global interaction stream, learns to detect and correct errors of the deformation field. We evaluate our two-stream approach for inpainting tasks, where experiments show that it improves both the propagation of features within a single frame as required for image inpainting, as well as their propagation from keyframes to target frames. Applied to video inpainting, our approach leads to 44% and 26% improvements in FID and LPIPS scores. Code at https://github.com/runwayml/guided-inpainting
Code (1)
Tasks
Image InpaintingVideo EditingVideo InpaintingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
Video Super-Resolution (VSR) aims to restore high-quality video frames from low-resolution (LR) estimates, yet most existing VSR approaches behave like black boxes at inference time: users cannot reliably correct unexpec…
Image Super-ResolutionVideo Super-ResolutionStyle TransferTemporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions
Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit fr…
Surgical phase recognitionOptical Flow EstimationRepresentation LearningMulti-Task LearningFrameONE: Hierarchical Motion Modeling for Universal Multi-View Echocardiographic Keyframe Detection
Accurate detection of end-systole (ES) and end-diastole (ED) frames is fundamental to echocardiographic assessment. Existing methods are typically developed in a view-specific manner, depend on auxiliary annotations or i…
Representation LearningMulti-Task LearningFlexible Motion In-betweening with Diffusion Models
Motion in-betweening, a fundamental task in character animation, consists of generating motion sequences that plausibly interpolate user-provided keyframe constraints. It has long been recognized as a labor-intensive and…
Imputationmotion in-betweeningThe Devil is in Temporal Token: High Quality Video Reasoning Segmentation
Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To o…
Reasoning SegmentationReferring Expression SegmentationReferring Video Object SegmentationSegmentation