Improving Semantic Segmentation through Spatio-Temporal Consistency Learned from Videos
We leverage unsupervised learning of depth, egomotion, and camera intrinsics to improve the performance of single-image semantic segmentation, by enforcing 3D-geometric and temporal consistency of segmentation masks across video frames. The predicted depth, egomotion, and camera intrinsics are used to provide an additional supervision signal to the segmentation model, significantly enhancing its quality, or, alternatively, reducing the number of labels the segmentation model needs. Our experiments were performed on the ScanNet dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
SegmentationSemantic SegmentationSimilar Papers 제목 키워드 기반
A spatio-temporal network for video semantic segmentation in surgical videos
Semantic segmentation in surgical videos has applications in intra-operative guidance, post-operative analytics and surgical education. Segmentation models need to provide accurate and consistent predictions since tempor…
DecoderSegmentationSemantic SegmentationVideo Semantic SegmentationSpatio-Temporal Attention for Consistent Video Semantic Segmentation in Automated Driving
Deep neural networks, especially transformer-based architectures, have achieved remarkable success in semantic segmentation for environmental perception. However, existing models process video frames independently, thus …
Video Semantic SegmentationComputational EfficiencyEvery Frame Counts: Joint Learning of Video Segmentation and Optical Flow
A major challenge for video semantic segmentation is the lack of labeled data. In most benchmark datasets, only one frame of a video clip is annotated, which makes most supervised methods fail to utilize information from…
Optical Flow EstimationSegmentationSemantic SegmentationVideo Segmentation+1RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation
Recently, state space models have demonstrated efficient video segmentation through linear-complexity state space compression. However, Video Semantic Segmentation (VSS) requires pixel-level spatiotemporal modeling capab…
Video Semantic SegmentationComputational EfficiencyVideo SegmentationSemantic Segmentation on VSPW Dataset through Masked Video Consistency
Pixel-level Video Understanding requires effectively integrating three-dimensional data in both spatial and temporal dimensions to learn accurate and stable semantic information from continuous frames. However, existing …
Semantic SegmentationVideo Understanding