Orientation-aware Semantic Segmentation on Icosahedron Spheres
We address semantic segmentation on omnidirectional images, to leverage a holistic understanding of the surrounding scene for applications like autonomous driving systems. For the spherical domain, several methods recently adopt an icosahedron mesh, but systems are typically rotation invariant or require significant memory and parameters, thus enabling execution only at very low resolutions. In our work, we propose an orientation-aware CNN framework for the icosahedron mesh. Our representation allows for fast network operations, as our design simplifies to standard network operations of classical CNNs, but under consideration of north-aligned kernel convolutions for features on the sphere. We implement our representation and demonstrate its memory efficiency up-to a level-8 resolution mesh (equivalent to 640 x 1024 equirectangular images). Finally, since our kernels operate on the tangent of the sphere, standard feature weights, pretrained on perspective data, can be directly transferred with only small need for weight refinement. In our evaluation our orientation-aware CNN becomes a new state of the art for the recent 2D3DS dataset, and our Omni-SYNTHIA version of SYNTHIA. Rotation invariant classification and segmentation tasks are additionally presented for comparison to prior art.
Code (1)
Tasks
Autonomous DrivingSemantic SegmentationSimilar Papers 제목 키워드 기반
Equivariant Networks for Pixelized Spheres
Pixelizations of Platonic solids such as the cube and icosahedron have been widely used to represent spherical data, from climate records to Cosmic Microwave Background maps. Platonic solids have well-known global symmet…
Semantic SegmentationSphereSR: 360° Image Super-Resolution with Arbitrary Projection via Continuous Spherical Image Representation
The 360{\deg}imaging has recently gained great attention; however, its angular resolution is relatively lower than that of a narrow field-of-view (FOV) perspective image as it is captured by using fisheye lenses with the…
ERPImage Super-ResolutionSuper-ResolutionSphereSR: 360deg Image Super-Resolution With Arbitrary Projection via Continuous Spherical Image Representation
The 360deg imaging has recently gained much attention; however, its angular resolution is relatively lower than that of a narrow field-of-view (FOV) perspective image as it is captured using a fisheye lens with the s…
ERPImage Super-ResolutionSuper-ResolutionLOGCAN++: Adaptive Local-global class-aware network for semantic segmentation of remote sensing imagery
Remote sensing images usually characterized by complex backgrounds, scale and orientation variations, and large intra-class variance. General semantic segmentation methods usually fail to fully investigate the above issu…
Image SegmentationSegmentationSegmentation Of Remote Sensing ImagerySemantic SegmentationElite360M: Efficient 360 Multi-task Learning via Bi-projection Fusion and Cross-task Collaboration
360 cameras capture the entire surrounding environment with a large FoV, exhibiting comprehensive visual information to directly infer the 3D structures, e.g., depth and surface normal, and semantic information simultane…
3D geometryERPMulti-Task LearningSemantic Segmentation+1