TFNet: Exploiting Temporal Cues for Fast and Accurate LiDAR Semantic Segmentation
LiDAR semantic segmentation plays a crucial role in enabling autonomous driving and robots to understand their surroundings accurately and robustly. A multitude of methods exist within this domain, including point-based, range-image-based, polar-coordinate-based, and hybrid strategies. Among these, range-image-based techniques have gained widespread adoption in practical applications due to their efficiency. However, they face a significant challenge known as the `many-to-one'' problem caused by the range image's limited horizontal and vertical angular resolution. As a result, around 20% of the 3D points can be occluded. In this paper, we present TFNet, a range-image-based LiDAR semantic segmentation method that utilizes temporal information to address this issue. Specifically, we incorporate a temporal fusion layer to extract useful information from previous scans and integrate it with the current scan. We then design a max-voting-based post-processing technique to correct false predictions, particularly those caused by the `many-to-one'' issue. We evaluated the approach on two benchmarks and demonstrated that the plug-in post-processing technique is generic and can be applied to various networks.
Code (0)
등록된 구현이 없습니다.
Tasks
Autonomous DrivingLIDAR Semantic SegmentationSemantic SegmentationSimilar Papers 제목 키워드 기반
Training-Time-Friendly Network for Real-Time Object Detection
Modern object detectors can rarely achieve short training time, fast inference speed, and high accuracy at the same time. To strike a balance among them, we propose the Training-Time-Friendly Network (TTFNet). In this wo…
Objectobject-detectionObject DetectionReal-Time Object DetectionEnd-to-End Neural Speech Coding for Real-Time Communications
Deep-learning based methods have shown their advantages in audio coding over traditional ones but limited attention has been paid on real-time communications (RTC). This paper proposes the TFNet, an end-to-end neural spe…
DecoderPacket Loss ConcealmentSpeech EnhancementMBTFNet: Multi-Band Temporal-Frequency Neural Network For Singing Voice Enhancement
A typical neural speech enhancement (SE) approach mainly handles speech and noise mixtures, which is not optimal for singing voice enhancement scenarios. Music source separation (MSS) models treat vocals and various acco…
Music Source SeparationSpeech EnhancementConfidence-guided Adaptive Gate and Dual Differential Enhancement for Video Salient Object Detection
Video salient object detection (VSOD) aims to locate and segment the most attractive object by exploiting both spatial cues and temporal cues hidden in video sequences. However, spatial and temporal cues are often unreli…
object-detectionObject DetectionOptical Flow EstimationSalient Object Detection+1EMGTFNet: Fuzzy Vision Transformer to decode Upperlimb sEMG signals for Hand Gestures Recognition
Myoelectric control is an area of electromyography of increasing interest nowadays, particularly in applications such as Hand Gesture Recognition (HGR) for bionic prostheses. Today's focus is on pattern recognition using…
Data AugmentationGesture RecognitionHand Gesture RecognitionHand-Gesture Recognition+1