paper-with-me

홈 › Papers

Leveraging Temporal Cues for Semi-Supervised Multi-View 3D Object Detection

2025-01-01 · CVPR 2025 1 · Jinhyung Park, Navyata Sanghvi, Hiroki Adachi, Yoshihisa Shibata, Shawn Hunt, Shinya Tanaka, Hironobu Fujiyoshi, Kris Kitani

While recent advancements in camera-based 3D object detection demonstrate remarkable performance, they require thousands or even millions of human-annotated frames. This requirement significantly inhibits their deployment in various locations and sensor configurations. To address this gap, we propose a performant semi-supervised framework that leverages unlabeled RGB-only driving sequences - data easily collected with cost-effective RGB cameras - to significantly improve temporal, camera-only 3D detectors. We observe that the standard semi-supervised pseudo-labeling paradigm underperforms in this temporal, camera-only setting due to poor 3D localization of pseudo-labels. To address this, we train a single 3D detector to handle RGB sequences both forward and backward in time, then ensemble both its forwards and backwards pseudo-labels for semi-supervised learning. We further improve the pseudo-label quality by leveraging 3D object tracking to infill missing detections and by eschewing simple confidence thresholding in favor of using the auxiliary 2D detection head to filter 3D predictions. Finally, to enable the backbone to learn directly from the unlabeled data itself, we introduce an object-query conditioned masked reconstruction objective. Our framework demonstrates remarkable performance improvement on large-scale autonomous driving datasets nuScenes and nuPlan.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object Detection3D Object TrackingAutonomous Drivingobject-detectionObject DetectionObject TrackingPseudo Label

Similar Papers 제목 키워드 기반

Semi-supervised Time Series Classification by Temporal Relation Prediction

2021-06-11 · ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2021 6 · Haoyi Fan, Fengbin Zhang, Ruidong Wang, Xunhua Huang 외

Semi-supervised learning (SSL) has proven to be a powerful algorithm in different domains by leveraging unlabeled data to mitigate the reliance on the tremendous annotated data. However, few efforts consider the underlyi…

ClassificationPredictionRelationRelation Prediction+4

Robust Semi-Supervised Monocular Depth Estimation with Reprojected Distances

2019-10-04 · Vitor Guizilini, Jie Li, Rares Ambrus, Sudeep Pillai 외

Dense depth estimation from a single image is a key problem in computer vision, with exciting applications in a multitude of robotic tasks. Initially viewed as a direct regression problem, requiring annotated labels as s…

Depth EstimationMonocular Depth Estimationvalid

Sample, Crop, Track: Self-Supervised Mobile 3D Object Detection for Urban Driving LiDAR

2022-09-21 · Sangyun Shin, Stuart Golodetz, Madhu Vankadari, Kaichen Zhou 외

Deep learning has led to great progress in the detection of mobile (i.e. movement-capable) objects in urban driving scenes in recent years. Supervised approaches typically require the annotation of large training sets; t…

3D Object DetectionObjectobject-detectionObject Detection+1

Disentangling spatio-temporal knowledge for weakly supervised object detection and segmentation in surgical video

2024-07-22 · Guiqiu Liao, Matjaz Jogan, Sai Koushik, Eric Eaton 외

Weakly supervised video object segmentation (WSVOS) enables the identification of segmentation maps without requiring an extensive training dataset of object masks, relying instead on coarse video labels indicating objec…

DisentanglementKnowledge DistillationObjectobject-detection+6

Semi-Supervised Video Salient Object Detection Using Pseudo-Labels

2019-08-12 · ICCV 2019 10 · Pengxiang Yan, Guanbin Li, Yuan Xie, Zhen Li 외

Deep learning-based video salient object detection has recently achieved great success with its performance significantly outperforming any other unsupervised methods. However, existing data-driven approaches heavily rel…

object-detectionRGB Salient Object DetectionSalient Object DetectionUnsupervised Video Object Segmentation+1