MRFP: Learning Generalizable Semantic Segmentation from Sim-2-Real with Multi-Resolution Feature Perturbation
Deep neural networks have shown exemplary performance on semantic scene understanding tasks on source domains, but due to the absence of style diversity during training, enhancing performance on unseen target domains using only single source domain data remains a challenging task. Generation of simulated data is a feasible alternative to retrieving large style-diverse real-world datasets as it is a cumbersome and budget-intensive process. However, the large domain-specfic inconsistencies between simulated and real-world data pose a significant generalization challenge in semantic segmentation. In this work, to alleviate this problem, we propose a novel MultiResolution Feature Perturbation (MRFP) technique to randomize domain-specific fine-grained features and perturb style of coarse features. Our experimental results on various urban-scene segmentation datasets clearly indicate that, along with the perturbation of style-information, perturbation of fine-feature components is paramount to learn domain invariant robust feature maps for semantic segmentation models. MRFP is a simple and computationally efficient, transferable module with no additional learnable parameters or objective functions, that helps state-of-the-art deep neural networks to learn robust domain invariant features for simulation-to-real semantic segmentation.
Code (1)
Tasks
2D Semantic SegmentationAutonomous DrivingAutonomous VehiclesDomain GeneralizationSemantic SegmentationStyle GeneralizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
RT-GS2: Real-Time Generalizable Semantic Segmentation for 3D Gaussian Representations of Radiance Fields
Gaussian Splatting has revolutionized the world of novel view synthesis by achieving high rendering performance in real-time. Recently, studies have focused on enriching these 3D representations with semantic information…
Novel View SynthesisSegmentationSemantic SegmentationA Class-wise Non-salient Region Generalized Framework for Video Semantic Segmentation
Video semantic segmentation (VSS) is beneficial for dealing with dynamic scenes due to the continuous property of the real-world environment. On the one hand, some methods alleviate the predicted inconsistent problem bet…
Domain GeneralizationSegmentationSemantic SegmentationVideo Semantic SegmentationTrueCity: Real and Simulated Urban Data for Cross-Domain 3D Scene Understanding
3D semantic scene understanding remains a long-standing challenge in the 3D computer vision community. One of the key issues pertains to limited real-world annotated data to facilitate generalizable models. The common pr…
Semantic SegmentationScene UnderstandingPoint CloudsOVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies
Open-vocabulary scene understanding using 3D Gaussian (3DGS) representations has garnered considerable attention. However, existing methods mostly lift knowledge from large 2D vision models into 3DGS on a scene-by-scene …
3DGS3D Semantic SegmentationOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic Segmentation+2Generalizable Medical Image Segmentation via Random Amplitude Mixup and Domain-Specific Image Restoration
For medical image analysis, segmentation models trained on one or several domains lack generalization ability to unseen domains due to discrepancies between different data acquisition policies. We argue that the degenera…
Image RestorationImage SegmentationMedical Image AnalysisMedical Image Segmentation+2