Predicting Future Instance Segmentation by Forecasting Convolutional Features
Anticipating future events is an important prerequisite towards intelligent behavior. Video forecasting has been studied as a proxy task towards this goal. Recent work has shown that to predict semantic segmentation of future frames, forecasting at the semantic level is more effective than forecasting RGB frames and then segmenting these. In this paper we consider the more challenging problem of future instance segmentation, which additionally segments out individual objects. To deal with a varying number of output labels per image, we develop a predictive model in the space of fixed-sized convolutional features of the Mask R-CNN instance segmentation model. We apply the "detection head'" of Mask R-CNN on the predicted features to produce the instance segmentation of future frames. Experiments show that this approach significantly improves over strong baselines based on optical flow and repurposed instance segmentation architectures.
Code (1)
Tasks
Instance SegmentationOptical Flow EstimationSegmentationSemantic SegmentationVideo ForecastingVideo PredictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Forecasting Future Instance Segmentation with Learned Optical Flow and Warping
For an autonomous vehicle it is essential to observe the ongoing dynamics of a scene and consequently predict imminent future scenarios to ensure safety to itself and others. This can be done using different sensors and …
Instance SegmentationOptical Flow EstimationSemantic SegmentationFinding Islands of Predictability in Action Forecasting
We address dense action forecasting: the problem of predicting future action sequence over long durations based on partial observation. Our key insight is that future action sequences are more accurately modeled with var…
Predicting Deeper into the Future of Semantic Segmentation
The ability to predict and therefore to anticipate the future is an important attribute of intelligence. It is also of utmost importance in real-time systems, e.g. in robotics or autonomous driving, which depend on visua…
AttributeAutonomous DrivingDecision MakingOptical Flow Estimation+3Dense Semantic Forecasting in Video by Joint Regression of Features and Feature Motion
Dense semantic forecasting anticipates future events in video by inferring pixel-level semantics of an unobserved future image. We present a novel approach that is applicable to various single-frame architectures and tas…
Future predictionPanoptic SegmentationregressionSegmentation+1CASNet: Common Attribute Support Network for image instance and panoptic segmentation
Instance segmentation and panoptic segmentation is being paid more and more attention in recent years. In comparison with bounding box based object detection and semantic segmentation, instance segmentation can provide m…
AttributeClusteringInstance Segmentationobject-detection+4