Motion-Appearance Interactive Encoding for Object Segmentation in Unconstrained Videos
We present a novel method of integrating motion and appearance cues for foreground object segmentation in unconstrained videos. Unlike conventional methods encoding motion and appearance patterns individually, our method puts particular emphasis on their mutual assistance. Specifically, we propose using an interactively constrained encoding (ICE) scheme to incorporate motion and appearance patterns into a graph that leads to a spatiotemporal energy optimization. The reason of utilizing ICE is that both motion and appearance cues for the same target share underlying correlative structure, thus can be exploited in a deeply collaborative manner. We perform ICE not only in the initialization but also in the refinement stage of a two-layer framework for object segmentation. This scheme allows our method to consistently capture structural patterns about object perceptions throughout the whole framework. Our method can be operated on superpixels instead of raw pixels to reduce the number of graph nodes by two orders of magnitude. Moreover, we propose to partially explore the multi-object localization problem with inter-occlusion by weighted bipartite graph matching. Comprehensive experiments on three benchmark datasets (i.e., SegTrack, MOViCS, and GaTech) demonstrate the effectiveness of our approach compared with extensive state-of-the-art methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Graph MatchingObjectObject LocalizationSemantic SegmentationSuperpixelsSimilar Papers 제목 키워드 기반
Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation
Recent mainstream unsupervised video object segmentation (UVOS) motion-appearance approaches use either the bi-encoder structure to separately encode motion and appearance features, or the uni-encoder structure for joint…
Optical Flow EstimationSalient Object DetectionSemantic SegmentationUnsupervised Video Object Segmentation+3MoSAM: Motion-Guided Segment Anything Model with Spatial-Temporal Memory Selection
The recent Segment Anything Model 2 (SAM2) has demonstrated exceptional capabilities in interactive object segmentation for both images and videos. However, as a foundational model on interactive segmentation, SAM2 perfo…
Instance SegmentationInteractive SegmentationObjectObject Tracking+5Interactive Segmentation of Radiance Fields
Radiance Fields (RF) are popular to represent casually-captured scenes for new view synthesis and several applications beyond it. Mixed reality on personal spaces needs understanding and manipulating scenes represented a…
Interactive SegmentationMixed RealitySegmentationSemantic SegmentationImproving Unsupervised Video Object Segmentation with Motion-Appearance Synergy
We present IMAS, a method that segments the primary objects in videos without manual annotation in training or inference. Previous methods in unsupervised video object segmentation (UVOS) have demonstrated the effectiven…
MisconceptionsObjectObject DiscoverySegmentation+4Gamifying Video Object Segmentation
Video object segmentation can be considered as one of the most challenging computer vision problems. Indeed, so far, no existing solution is able to effectively deal with the peculiarities of real-world videos, especiall…
Interactive Video Object SegmentationObjectSegmentationSemantic Segmentation+2