DyStaB: Unsupervised Object Segmentation via Dynamic-Static Bootstrapping
We describe an unsupervised method to detect and segment portions of images of live scenes that, at some point in time, are seen moving as a coherent whole, which we refer to as objects. Our method first partitions the motion field by minimizing the mutual information between segments. Then, it uses the segments to learn object models that can be used for detection in a static image. Static and dynamic models are represented by deep neural networks trained jointly in a bootstrapping strategy, which enables extrapolation to previously unseen objects. While the training process requires motion, the resulting object segmentation network can be used on either static images or videos at inference time. As the volume of seen videos grows, more and more objects are seen moving, priming their detection, which then serves as a regularizer for new objects, turning our method into unsupervised continual learning to segment objects. Our models are compared to the state of the art in both video object segmentation and salient object detection. In the six benchmark datasets tested, our models compare favorably even to those using pixel-level supervision, despite requiring no manual annotation.
Code (0)
등록된 구현이 없습니다.
Tasks
Continual LearningObjectobject-detectionObject DetectionRGB Salient Object DetectionSalient Object DetectionSemantic SegmentationUnsupervised Object SegmentationUnsupervised Video Object SegmentationVideo Object SegmentationVideo Semantic SegmentationSimilar Papers 제목 키워드 기반
Dynamic semantic VSLAM with known and unknown objects
Traditional Visual Simultaneous Localization and Mapping (VSLAM) systems assume a static environment, which makes them ineffective in highly dynamic settings. To overcome this, many approaches integrate semantic informat…
Optical Flow EstimationSimultaneous Localization and MappingInstance Embedding Transfer to Unsupervised Video Object Segmentation
We propose a method for unsupervised video object segmentation by transferring the knowledge encapsulated in image-based instance embedding networks. The instance embedding network produces an embedding vector for each p…
ObjectOptical Flow EstimationSegmentationSemantic Segmentation+34DContrast: Contrastive Learning with Dynamic Correspondences for 3D Scene Understanding
We present a new approach to instill 4D dynamic object priors into learned 3D representations by unsupervised pre-training. We observe that dynamic movement of an object through an environment provides important cues abo…
3D Instance Segmentation3D Semantic SegmentationContrastive LearningData Augmentation+9Learning Unsupervised Video Object Segmentation Through Visual Attention
This paper conducts a systematic study on the role of visual attention in Unsupervised Video Object Segmentation (UVOS) tasks. By elaborately annotating three popular video segmentation datasets (DAVIS, Youtube-Objects a…
ObjectSegmentationSemantic SegmentationUnsupervised Video Object Segmentation+3Semantically Coherent Co-Segmentation and Reconstruction of Dynamic Scenes
In this paper we propose a framework for spatially and temporally coherent semantic co-segmentation and reconstruction of complex dynamic scenes from multiple static or moving cameras. Semantic co-segmentation exploits t…
3D ReconstructionSegmentation