paper-with-me

홈 › Papers

ActiveAnno3D -- An Active Learning Framework for Multi-Modal 3D Object Detection

2024-02-05 · Ahmed Ghita, Bjørk Antoniussen, Walter Zimmer, Ross Greer, Christian Creß, Andreas Møgelmose, Mohan M. Trivedi, Alois C. Knoll

The curation of large-scale datasets is still costly and requires much time and resources. Data is often manually labeled, and the challenge of creating high-quality datasets remains. In this work, we fill the research gap using active learning for multi-modal 3D object detection. We propose ActiveAnno3D, an active learning framework to select data samples for labeling that are of maximum informativeness for training. We explore various continuous training methods and integrate the most efficient method regarding computational demand and detection performance. Furthermore, we perform extensive experiments and ablation studies with BEVFusion and PV-RCNN on the nuScenes and TUM Traffic Intersection dataset. We show that we can achieve almost the same performance with PV-RCNN and the entropy-based query strategy when using only half of the training data (77.25 mAP compared to 83.50 mAP) of the TUM Traffic Intersection dataset. BEVFusion achieved an mAP of 64.31 when using half of the training data and 75.0 mAP when using the complete nuScenes dataset. We integrate our active learning framework into the proAnno labeling tool to enable AI-assisted data selection and labeling and minimize the labeling costs. Finally, we provide code, weights, and visualization results on our website: https://active3d-framework.github.io/active3d-framework.

📄 PDF Abstract BibTeX arXiv:2402.03235

Code (1)

walzimmer/3d-bat

Tasks

3D Object DetectionActive LearningInformativenessobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Batch Normalization 설명 없음
TUM 설명 없음

Similar Papers 제목 키워드 기반

ActiveAnno: General-Purpose Document-Level Annotation Tool with Active Learning Integration

2021-06-01 · NAACL 2021 4 · Max Wiechmann, Seid Muhie Yimam, Chris Biemann

ActiveAnno is an annotation tool focused on document-level annotation tasks developed both for industry and research settings. It is designed to be a general-purpose tool with a wide variety of use cases. It features a m…

Active Learning

UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation

2021-09-13 · Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 외

With the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and visual modalities to o…

Abstractive Text SummarizationDecoderImage CaptioningKnowledge Distillation+1

Multi-Modal Summary Generation using Multi-Objective Optimization

2020-05-19 · Anubhav Jangra, Sriparna Saha, Adam Jatowt, Mohammad Hasanuzzaman

Significant development of communication technology over the past few years has motivated research in multi-modal summarization techniques. A majority of the previous works on multi-modal summarization focus on text and …

Active Speaker Detection as a Multi-Objective Optimization with Uncertainty-based Multimodal Fusion

2021-06-07 · Baptiste Pouthier, Laurent Pilati, Leela K. Gudupudi, Charles Bouveyron 외

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovi…

Active Speaker DetectionAudio-Visual Active Speaker Detection

MARS: Multimodal Active Robotic Sensing for Articulated Characterization

2024-07-01 · Hongliang Zeng, Ping Zhang, Chengjiong Wu, Jiahua Wang 외

Precise perception of articulated objects is vital for empowering service robots. Recent studies mainly focus on point cloud, a single-modal approach, often neglecting vital texture and lighting details and assuming idea…

parameter estimation