Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation
Location and appearance are the key cues for video object segmentation. Many sources such as RGB, depth, optical flow and static saliency can provide useful information about the objects. However, existing approaches only utilize the RGB or RGB and optical flow. In this paper, we propose a novel multi-source fusion network for zero-shot video object segmentation. With the help of interoceptive spatial attention module (ISAM), spatial importance of each source is highlighted. Furthermore, we design a feature purification module (FPM) to filter the inter-source incompatible features. By the ISAM and FPM, the multi-source features are effectively fused. In addition, we put forward an automatic predictor selection network (APS) to select the better prediction of either the static saliency predictor or the moving object predictor in order to prevent over-reliance on the failed results caused by low-quality optical flow maps. Extensive experiments on three challenging public benchmarks (i.e. DAVIS$_{16}$, Youtube-Objects and FBMS) show that the proposed model achieves compelling performance against the state-of-the-arts. The source code will be publicly available at \textcolor{red}{\url{https://github.com/Xiaoqi-Zhao-DLUT/Multi-Source-APS-ZVOS}}.
Code (1)
Tasks
Depth EstimationObjectSalient Object DetectionUnsupervised Video Object SegmentationVideo Object SegmentationZero-Shot Video Object SegmentationSimilar Papers 제목 키워드 기반
Diffusion-Driven High-Dimensional Variable Selection
Variable selection for high-dimensional, highly correlated data has long been a challenging problem, often yielding unstable and unreliable models. We propose a resample-aggregate framework that exploits diffusion models…
Transfer LearningData AugmentationControllable Accent Normalization via Discrete Diffusion
Existing accent normalization methods do not typically offer control over accent strength, yet many applications-such as language learning and dubbing-require tunable accent retention. We propose DLM-AN, a controllable a…
Adaptive Multi-source Predictor for Zero-shot Video Object Segmentation
Static and moving objects often occur in real-life videos. Most video object segmentation methods only focus on extracting and exploiting motion cues to perceive moving objects. Once faced with the frames of static objec…
ObjectOptical Flow EstimationSemantic SegmentationUnsupervised Video Object Segmentation+3Joint Manifold Diffusion for Combining Predictions on Decoupled Observations
We present a new predictor combination algorithm that improves a given task predictor based on potentially relevant reference predictors. Existing approaches are limited in that, to discover the underlying task dependenc…
Source Data Selection for Brain-Computer Interfaces based on Simple Features
This paper demonstrates that simple features available during the calibration of a brain-computer interface can be utilized for source data selection to improve the performance of the brain-computer interface for a new t…
Brain Computer InterfaceMotor ImageryTransfer Learning