Future Segmentation Using 3D Structure
Predicting the future to anticipate the outcome of events and actions is a critical attribute of autonomous agents; particularly for agents which must rely heavily on real time visual data for decision making. Working towards this capability, we address the task of predicting future frame segmentation from a stream of monocular video by leveraging the 3D structure of the scene. Our framework is based on learnable sub-modules capable of predicting pixel-wise scene semantic labels, depth, and camera ego-motion of adjacent frames. We further propose a recurrent neural network based model capable of predicting future ego-motion trajectory as a function of a series of past ego-motion steps. Ultimately, we observe that leveraging 3D structure in the model facilitates successful prediction, achieving state of the art accuracy in future semantic segmentation.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeDecision MakingSegmentationSemantic SegmentationSimilar Papers 제목 키워드 기반
Unsupervised Video Prediction from a Single Frame by Estimating 3D Dynamic Scene Structure
Our goal in this work is to generate realistic videos given just one initial frame as input. Existing unsupervised approaches to this task do not consider the fact that a video typically shows a 3D environment, and that …
Motion SegmentationSegmentationVideo PredictionHuman Treelike Tubular Structure Segmentation: A Comprehensive Review and Future Perspectives
Various structures in human physiology follow a treelike morphology, which often expresses complexity at very fine scales. Examples of such structures are intrathoracic airways, retinal blood vessels, and hepatic blood v…
Computed Tomography (CT)PrognosisPredictive Feature Learning for Future Segmentation Prediction
Future segmentation prediction aims to predict the segmentation masks for unobserved future frames. Most existing works addressed it by directly predicting the intermediate features extracted by existing segmentation…
PredictionSegmentationSegment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
Video Object Segmentation and Tracking (VOST) presents a complex yet critical challenge in computer vision, requiring robust integration of segmentation and tracking across temporally dynamic frames. Traditional methods …
Video Object SegmentationComputational EfficiencyDomain GeneralizationExtending Pretrained Segmentation Networks with Additional Anatomical Structures
Comprehensive surgical planning require complex patient-specific anatomical models. For instance, functional muskuloskeletal simulations necessitate all relevant structures to be segmented, which could be performed in re…
class-incremental learningClass Incremental LearningIncremental LearningSegmentation