Treating Motion as Option to Reduce Motion Dependency in Unsupervised Video Object Segmentation
Unsupervised video object segmentation (VOS) aims to detect the most salient object in a video sequence at the pixel level. In unsupervised VOS, most state-of-the-art methods leverage motion cues obtained from optical flow maps in addition to appearance cues to exploit the property that salient objects usually have distinctive movements compared to the background. However, as they are overly dependent on motion cues, which may be unreliable in some cases, they cannot achieve stable prediction. To reduce this motion dependency of existing two-stream VOS methods, we propose a novel motion-as-option network that optionally utilizes motion cues. Additionally, to fully exploit the property of the proposed network that motion is not always required, we introduce a collaborative network learning strategy. On all the public benchmark datasets, our proposed network affords state-of-the-art performance with real-time inference speed.
Code (2)
Tasks
Optical Flow EstimationSemantic SegmentationUnsupervised Video Object SegmentationVideo Object SegmentationVideo Semantic SegmentationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Treating Motion as Option with Output Selection for Unsupervised Video Object Segmentation
Unsupervised video object segmentation (VOS) is a task that aims to detect the most salient object in a video without external guidance about the object. To leverage the property that salient objects usually have distinc…
ObjectOptical Flow EstimationSemantic SegmentationUnsupervised Video Object Segmentation+2The sub-fractional CEV model
The sub-fractional Brownian motion (sfBm) is a stochastic process, characterized by non-stationarity in their increments and long-range dependency, considered as an intermediate step between the standard Brownian motion …
modelSynthesizing Long-Term Human Motions with Diffusion Models via Coherent Sampling
Text-to-motion generation has gained increasing attention, but most existing methods are limited to generating short-term motions that correspond to a single sentence describing a single action. However, when a text stre…
Motion GenerationSentenceLearning Trajectory Dependencies for Human Motion Prediction
Human motion prediction, i.e., forecasting future body poses given observed pose sequence, has typically been tackled with recurrent neural networks (RNNs). However, as evidenced by prior work, the resulted RNN models su…
Human motion predictionHuman Pose Forecastingmotion predictionMulti-Person Pose forecasting+1StableFace: Analyzing and Improving Motion Stability for Talking Face Generation
While previous speech-driven talking face generation methods have made significant progress in improving the visual quality and lip-sync quality of the synthesized videos, they pay less attention to lip motion jitters wh…
Face GenerationTalking Face GenerationVideo Generation