paper-with-me

홈 › Papers

V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D Annotations

2024-12-16 · Jin-Cheng Jhang, Tao Tu, Fu-En Wang, Ke Zhang, Min Sun, Cheng-Hao Kuo

The field of indoor monocular 3D object detection is gaining significant attention, fueled by the increasing demand in VR/AR and robotic applications. However, its advancement is impeded by the limited availability and diversity of 3D training data, owing to the labor-intensive nature of 3D data collection and annotation processes. In this paper, we present V-MIND (Versatile Monocular INdoor Detector), which enhances the performance of indoor 3D detectors across a diverse set of object classes by harnessing publicly available large-scale 2D datasets. By leveraging well-established monocular depth estimation techniques and camera intrinsic predictors, we can generate 3D training data by converting large-scale 2D images into 3D point clouds and subsequently deriving pseudo 3D bounding boxes. To mitigate distance errors inherent in the converted point clouds, we introduce a novel 3D self-calibration loss for refining the pseudo 3D bounding boxes during training. Additionally, we propose a novel ambiguity loss to address the ambiguity that arises when introducing new classes from 2D datasets. Finally, through joint training with existing 3D datasets and pseudo 3D bounding boxes derived from 2D datasets, V-MIND achieves state-of-the-art object detection performance across a wide range of classes on the Omni3D indoor dataset.

📄 PDF Abstract BibTeX arXiv:2412.11412

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionDepth EstimationMonocular 3D Object DetectionMonocular Depth Estimationobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

ORB-SLAM: a Versatile and Accurate Monocular SLAM System

2015-02-03 · Raul Mur-Artal, J. M. M. Montiel, Juan D. Tardos

This paper presents ORB-SLAM, a feature-based monocular SLAM system that operates in real time, in small and large, indoor and outdoor environments. The system is robust to severe motion clutter, allows wide baseline loo…

Simultaneous Localization and Mapping

REMIND: RE-Identification with Memory for INDoor Navigation

2026-07-10 · Pablo Diaz-Pereda, Alejandro Rodriguez-Ramos, David Perez-Saura, Pascual Campoy arxiv

Mobile robots operating indoors must re-identify previously observed objects after long temporal gaps, significant viewpoint changes, and severe illumination variations. This remains a challenging problem: multi-object t…

Video Object SegmentationVehicle Re-IdentificationMulti-Object Tracking

NAVREN-RL: Learning to fly in real environment via end-to-end deep reinforcement learning using monocular images

2018-07-22 · Malik Aqeel Anwar, Arijit Raychowdhury

We present NAVREN-RL, an approach to NAVigate an unmanned aerial vehicle in an indoor Real ENvironment via end-to-end reinforcement learning RL. A suitable reward function is designed keeping in mind the cost and weight …

Deep Reinforcement LearningNavigatereinforcement-learningReinforcement Learning+1

Versatile Depth Estimator Based on Common Relative Depth Estimation and Camera-Specific Relative-to-Metric Depth Conversion

2023-03-20 · Jinyoung Jun, Jae-Han Lee, Chang-Su Kim

A typical monocular depth estimator is trained for a single camera, so its performance drops severely on images taken with different cameras. To address this issue, we propose a versatile depth estimator (VDE), composed …

Depth Estimation

SING3R-SLAM: Submap-based Indoor Monocular Gaussian SLAM with 3D Reconstruction Priors

2025-11-21 · Kunyi Li, Michael Niemeyer, Sen Wang, Stefano Gasperini 외 arxiv

Recent advances in dense 3D reconstruction have demonstrated strong capability in accurately capturing local geometry. However, extending these methods to incremental global reconstruction, as required in SLAM systems, r…

3D ReconstructionPose Estimation