Automatic Video Object Segmentation via Motion-Appearance-Stream Fusion and Instance-aware Segmentation
This paper presents a method for automatic video object segmentation based on the fusion of motion stream, appearance stream, and instance-aware segmentation. The proposed scheme consists of a two-stream fusion network and an instance segmentation network. The two-stream fusion network again consists of motion and appearance stream networks, which extract long-term temporal and spatial information, respectively. Unlike the existing two-stream fusion methods, the proposed fusion network blends the two streams at the original resolution for obtaining accurate segmentation boundary. We develop a recurrent bidirectional multiscale structure with skip connection for the stream fusion network to extract long-term temporal information. Also, the multiscale structure enables to obtain the original resolution features at the end of the network. As a result of two-stream fusion, we have a pixel-level probabilistic segmentation map, which has higher values at the pixels belonging to the foreground object. By combining the probability of foreground map and objectness score of instance segmentation mask, we finally obtain foreground segmentation results for video sequences without any user intervention, i.e., we achieve successful automatic video segmentation. The proposed structure shows a state-of-the-art performance for automatic video object segmentation task, and also achieves near semi-supervised performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Foreground SegmentationInstance SegmentationObjectSegmentationSemantic SegmentationVideo Object SegmentationVideo SegmentationVideo Semantic SegmentationSimilar Papers 제목 키워드 기반
FusionSeg: Learning to Combine Motion and Appearance for Fully Automatic Segmentation of Generic Objects in Videos
We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in …
SegmentationStructured PredictionVideo SegmentationVideo Semantic SegmentationFusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos
We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in …
SegmentationStructured PredictionUnsupervised Video Object SegmentationVideo Segmentation+1The Emergence of Objectness: Learning Zero-Shot Segmentation from Videos
Humans can easily segment moving objects without knowing what they are. That objectness could emerge from continuous visual observations motivates us to model grouping and movement concurrently from unlabeled videos. Our…
Contrastive LearningImage SegmentationSegmentationSemantic Segmentation+4CamoSAM2: Motion-Appearance Induced Auto-Refining Prompts for Video Camouflaged Object Detection
The Segment Anything Model 2 (SAM2), a prompt-guided video foundation model, has remarkably performed in video object segmentation, drawing significant attention in the community. Due to the high similarity between camou…
Camouflaged Object Segmentationobject-detectionObject DetectionSemantic Segmentation+2Improving Unsupervised Video Object Segmentation with Motion-Appearance Synergy
We present IMAS, a method that segments the primary objects in videos without manual annotation in training or inference. Previous methods in unsupervised video object segmentation (UVOS) have demonstrated the effectiven…
MisconceptionsObjectObject DiscoverySegmentation+4