Two-Stream Networks for Object Segmentation in Videos
Existing matching-based approaches perform video object segmentation (VOS) via retrieving support features from a pixel-level memory, while some pixels may suffer from lack of correspondence in the memory (i.e., unseen), which inevitably limits their segmentation performance. In this paper, we present a Two-Stream Network (TSN). Our TSN includes (i) a pixel stream with a conventional pixel-level memory, to segment the seen pixels based on their pixellevel memory retrieval. (ii) an instance stream for the unseen pixels, where a holistic understanding of the instance is obtained with dynamic segmentation heads conditioned on the features of the target instance. (iii) a pixel division module generating a routing map, with which output embeddings of the two streams are fused together. The compact instance stream effectively improves the segmentation accuracy of the unseen pixels, while fusing two streams with the adaptive routing map leads to an overall performance boost. Through extensive experiments, we demonstrate the effectiveness of our proposed TSN, and we also report state-of-the-art performance of 86.1% on YouTube-VOS 2018 and 87.5% on the DAVIS-2017 validation split.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectRetrievalSegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic SegmentationVocal Bursts Valence PredictionSimilar Papers 제목 키워드 기반
Hierarchical Deep Co-segmentation of Primary Objects in Aerial Videos
Primary object segmentation plays an important role in understanding videos generated by unmanned aerial vehicles. In this paper, we propose a large-scale dataset with 500 aerial videos and manually annotated primary obj…
SegmentationSemantic SegmentationFusionSeg: Learning to Combine Motion and Appearance for Fully Automatic Segmentation of Generic Objects in Videos
We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in …
SegmentationStructured PredictionVideo SegmentationVideo Semantic SegmentationTemporally Object-based Video Co-Segmentation
In this paper, we propose an unsupervised video object co-segmentation framework based on the primary object proposals to extract the common foreground object(s) from a given video set. In addition to the objectness attr…
ObjectSegmentationFusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos
We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in …
SegmentationStructured PredictionUnsupervised Video Object SegmentationVideo Segmentation+13D-Aware Instance Segmentation and Tracking in Egocentric Videos
Egocentric videos present unique challenges for 3D scene understanding due to rapid camera motion, frequent object occlusions, and limited object visibility. This paper introduces a novel approach to instance segmentatio…
3D Object ReconstructionInstance SegmentationObjectObject Reconstruction+5