Separable Structure Modeling for Semi-supervised Video Object Segmentation
In this paper, we propose a separable structure modeling approach for semi-supervised video object segmentation. Unlike most existing methods which preclude the semantically structural information of target objects, our method not only captures pixel-level similarity relationships between the reference and target frames but also reveals the separable structure of the specified objects in target frames. Specifically, we first compute a pixel-wise similarity matrix by using representations of reference and target pixels and then select top-rank reference pixels for target pixel classification. According to the prior knowledge from these top-rank reference pixels, we further appoint the representative target pixels for object structure modeling. Particularly, in the structure modeling branch, we extract the shared and individual features that can well represent the whole object and its components, respectively. Moreover, the proposed method is a fast algorithm without online fine-tuning and any post-processing. We conduct extensive experiments and ablation studies on the DAVIS-16, DAVIS-17, and YouTube-VOS datasets, and experimental results on three widely-used datasets demonstrate that our method achieves superior performance, compared with state-of-the-art semi-supervised video object segmentation approaches in terms of speed and accuracy.
Code (1)
Tasks
ObjectOne-shot visual object segmentationSemi-Supervised Video Object SegmentationVideo Object SegmentationVideo Semantic SegmentationSimilar Papers 제목 키워드 기반
One-Class Semi-Supervised Learning: Detecting Linearly Separable Class by its Mean
In this paper, we presented a novel semi-supervised one-class classification algorithm which assumes that class is linearly separable from other elements. We proved theoretically that class is linearly separable if and o…
General ClassificationOne-Class ClassificationEnd-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
We study pseudo-labeling for the semi-supervised training of ResNet, Time-Depth Separable ConvNets, and Transformers for speech recognition, with either CTC or Seq2Seq loss functions. We perform experiments on the standa…
Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSpatiotemporal Classification with limited labels using Constrained Clustering for large datasets
Creating separable representations via representation learning and clustering is critical in analyzing large unstructured datasets with only a few labels. Separable representations can lead to supervised models with bett…
ClusteringConstrained ClusteringRepresentation LearningSemi-Supervised Domain Adaptation for Weakly Labeled Semantic Video Object Segmentation
Deep convolutional neural networks (CNNs) have been immensely successful in many high-level computer vision tasks given large labeled datasets. However, for video semantic object segmentation, a domain where labels are s…
Domain AdaptationSegmentationSemantic SegmentationSemi-supervised Domain Adaptation+3Recurrent Ladder Networks
We propose a recurrent extension of the Ladder networks whose structure is motivated by the inference required in hierarchical latent variable models. We demonstrate that the recurrent Ladder is able to handle a wide var…
Music Modeling