SurfaceAug: Closing the Gap in Multimodal Ground Truth Sampling
Despite recent advances in both model architectures and data augmentation, multimodal object detectors still barely outperform their LiDAR-only counterparts. This shortcoming has been attributed to a lack of sufficiently powerful multimodal data augmentation. To address this, we present SurfaceAug, a novel ground truth sampling algorithm. SurfaceAug pastes objects by resampling both images and point clouds, enabling object-level transformations in both modalities. We evaluate our algorithm by training a multimodal detector on KITTI and compare its performance to previous works. We show experimentally that SurfaceAug outperforms existing methods on car detection tasks and establishes a new state of the art for multimodal ground truth sampling.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationObjectSimilar Papers 제목 키워드 기반
Object sieving and morphological closing to reduce false detections in wide-area aerial imagery
For object detection in wide-area aerial imagery, post-processing is usually needed to reduce false detections. We propose a two-stage post-processing scheme which comprises an area-thresholding sieving process and a mor…
Objectobject-detectionObject DetectionFalse Positive Sampling-based Data Augmentation for Enhanced 3D Object Detection Accuracy
Recent studies have focused on enhancing the performance of 3D object detection models. Among various approaches, ground-truth sampling has been proposed as an augmentation technique to address the challenges posed by li…
3D Object DetectionData AugmentationObjectobject-detection+1FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
The accurate trust assessment of multimodal large language models (MLLMs) generated predictions, which can enable selective prediction and improve user confidence, is challenging due to the diverse multi-modal input para…
How Can I Publish My LLM Benchmark Without Giving the True Answers Away?
Publishing a large language model (LLM) benchmark on the Internet risks contaminating future LLMs: the benchmark may be unintentionally (or intentionally) used to train or select a model. A common mitigation is to keep t…
Large Language ModelGeneration of annotated multimodal ground truth datasets for abdominal medical image registration
Sparsity of annotated data is a major limitation in medical image processing tasks such as registration. Registered multimodal image data are essential for the diagnosis of medical conditions and the success of intervent…
Computed Tomography (CT)Image RegistrationImage SegmentationMedical Image Registration+1