StixelNExT: Toward Monocular Low-Weight Perception for Object Segmentation and Free Space Detection
In this work, we present a novel approach for general object segmentation from a monocular image, eliminating the need for manually labeled training data and enabling rapid, straightforward training and adaptation with minimal data. Our model initially learns from LiDAR during the training process, which is subsequently removed from the system, allowing it to function solely on monocular imagery. This study leverages the concept of the Stixel-World to recognize a medium level representation of its surroundings. Our network directly predicts a 2D multi-layer Stixel-World and is capable of recognizing and locating multiple, superimposed objects within an image. Due to the scarcity of comparable works, we have divided the capabilities into modules and present a free space detection in our experiments section. Furthermore, we introduce an improved method for generating Stixels from LiDAR data, which we use as ground truth for our network.
Code (1)
Tasks
Semantic SegmentationSimilar Papers 제목 키워드 기반
StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception
This paper presents StixelNExT++, a novel approach to scene representation for monocular perception systems. Building on the established Stixel representation, our method infers 3D Stixels and enhances object segmentatio…
Object SegmentationScene SegmentationSingle-Eye View: Monocular Real-time Perception Package for Autonomous Driving
Amidst the rapid advancement of camera-based autonomous driving technology, effectiveness is often prioritized with limited attention to computational efficiency. To address this issue, this paper introduces LRHPerceptio…
Computational EfficiencyTrajectory PredictionAutonomous DrivingRoad SegmentationMonocular Robot Navigation with Self-Supervised Pretrained Vision Transformers
In this work, we consider the problem of learning a perception model for monocular robot navigation using few annotated images. Using a Vision Transformer (ViT) pretrained with a label-free self-supervised method, we suc…
CPUImage SegmentationRobot NavigationSegmentation+1Monocular Depth Estimation and Segmentation for Transparent Object with Iterative Semantic and Geometric Fusion
Transparent object perception is indispensable for numerous robotic tasks. However, accurately segmenting and estimating the depth of transparent objects remain challenging due to complex optical properties. Existing met…
Depth EstimationMonocular Depth EstimationTransparent objectsM2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
Deploying real-time spatial perception on edge devices requires efficient multi-task models that leverage complementary task information while minimizing computational overhead. This paper introduces Multi-Mono-Hydra (M2…
Computational EfficiencySemantic SegmentationMulti-Task Learning