Training Weakly Supervised Video Frame Interpolation With Events
Event-based video frame interpolation is promising as event cameras capture dense motion signals that can greatly facilitate motion-aware synthesis. However, training existing frameworks for this task requires high frame-rate videos with synchronized events, posing challenges to collect real training data. In this work we show event-based frame interpolation can be trained without the need of high framerate videos. This is achieved via a novel weakly supervised framework that 1) corrects image appearance by extracting complementary information from events and 2) supplants motion dynamics modeling with attention mechanisms. For the latter we propose subpixel attention learning, which supports searching high-resolution correspondence efficiently on low-resolution feature grid. Though trained on low frame-rate videos, our framework outperforms existing models trained with full high frame-rate videos (and events) on both GoPro dataset and a new real event-based dataset. Codes, models and dataset will be made available at: https://github.com/YU-Zhiyang/WEVI.
Code (1)
Tasks
Video Frame InterpolationSimilar Papers 제목 키워드 기반
Video Shadow Detection via Spatio-Temporal Interpolation Consistency Training
It is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal in…
Shadow DetectionSpatial InterpolationVideo Shadow DetectionTowards Weakly Supervised End-to-end Learning for Long-video Action Recognition
Developing end-to-end action recognition models on long videos is fundamental and crucial for long-video action understanding. Due to the unaffordable cost of end-to-end training on the whole long videos, existing works …
Action ClassificationAction RecognitionAction SegmentationAction Understanding+5Learning a Weakly-Supervised Video Actor-Action Segmentation Model with a Wise Selection
We address weakly-supervised video actor-action segmentation (VAAS), which extends general video object segmentation (VOS) to additionally consider action labels of the actors. The most successful methods on VOS synthesi…
Action SegmentationSegmentationSemantic SegmentationVideo Object Segmentation+1Unsupervised Video Interpolation Using Cycle Consistency
Learning to synthesize high frame rate videos via interpolation requires large quantities of high frame rate training videos, which, however, are scarce, especially at high resolutions. Here, we propose unsupervised tech…
TripletVideo Frame InterpolationWeakly Supervised Instance Segmentation for Videos with Temporal Mask Consistency
Weakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) parti…
Instance SegmentationRelation NetworkSegmentationSemantic Segmentation+1