paper-with-me

홈 › Papers

Towards Unified Keyframe Propagation Models

2022-05-19 · Patrick Esser, Peter Michael, Soumyadip Sengupta

Many video editing tasks such as rotoscoping or object removal require the propagation of context across frames. While transformers and other attention-based approaches that aggregate features globally have demonstrated great success at propagating object masks from keyframes to the whole video, they struggle to propagate high-frequency details such as textures faithfully. We hypothesize that this is due to an inherent bias of global attention towards low-frequency features. To overcome this limitation, we present a two-stream approach, where high-frequency features interact locally and low-frequency features interact globally. The global interaction stream remains robust in difficult situations such as large camera motions, where explicit alignment fails. The local interaction stream propagates high-frequency details through deformable feature aggregation and, informed by the global interaction stream, learns to detect and correct errors of the deformation field. We evaluate our two-stream approach for inpainting tasks, where experiments show that it improves both the propagation of features within a single frame as required for image inpainting, as well as their propagation from keyframes to target frames. Applied to video inpainting, our approach leads to 44% and 26% improvements in FID and LPIPS scores. Code at https://github.com/runwayml/guided-inpainting

📄 PDF Abstract BibTeX arXiv:2205.09731

Code (1)

runwayml/guided-inpainting 공식 구현 pytorch

Tasks

Image InpaintingVideo EditingVideo Inpainting

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation

2026-03-17 · Jiongze Yu, Xiangbo Gao, Pooja Verlani, Akshay Gadde 외 arxiv

Video Super-Resolution (VSR) aims to restore high-quality video frames from low-resolution (LR) estimates, yet most existing VSR approaches behave like black boxes at inference time: users cannot reliably correct unexpec…

Image Super-ResolutionVideo Super-ResolutionStyle Transfer

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

2026-06-25 · Garam Kim, Juyoun Park arxiv

Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit fr…

Surgical phase recognitionOptical Flow EstimationRepresentation LearningMulti-Task Learning

FrameONE: Hierarchical Motion Modeling for Universal Multi-View Echocardiographic Keyframe Detection

2026-07-01 · Rusi Chen, Yuhao Huang, Hongyuan Zhang, Chao Tian 외 arxiv

Accurate detection of end-systole (ES) and end-diastole (ED) frames is fundamental to echocardiographic assessment. Existing methods are typically developed in a view-specific manner, depend on auxiliary annotations or i…

Representation LearningMulti-Task Learning

Flexible Motion In-betweening with Diffusion Models

2024-05-17 · Setareh Cohan, Guy Tevet, Daniele Reda, Xue Bin Peng 외

Motion in-betweening, a fundamental task in character animation, consists of generating motion sequences that plausibly interpolate user-provided keyframe constraints. It has long been recognized as a labor-intensive and…

Imputationmotion in-betweening

The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

2025-01-15 · CVPR 2025 1 · Sitong Gong, Yunzhi Zhuge, Lu Zhang, Zongxin Yang 외

Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To o…

Reasoning SegmentationReferring Expression SegmentationReferring Video Object SegmentationSegmentation