paper-with-me

Papers

Automated Multimodal Data Annotation via Calibration With Indoor Positioning System

2023-12-06 · Ryan Rubel, Andrew Dudash, Mohammad Goli, James O'Hara, Karl Wunderlich

Learned object detection methods based on fusion of LiDAR and camera data require labeled training samples, but niche applications, such as warehouse robotics or automated infrastructure, require semantic classes not available in large existing datasets. Therefore, to facilitate the rapid creation of multimodal object detection datasets and alleviate the burden of human labeling, we propose a novel automated annotation pipeline. Our method uses an indoor positioning system (IPS) to produce accurate detection labels for both point clouds and images and eliminates manual annotation entirely. In an experiment, the system annotates objects of interest 261.8 times faster than a human baseline and speeds up end-to-end dataset creation by 61.5%.

📄 PDF Abstract BibTeX arXiv:2312.03608

Code (0)

등록된 구현이 없습니다.

Tasks

Objectobject-detectionObject Detection

Similar Papers 제목 키워드 기반

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles

2025-09-10 · Eric Slyman, Mehrab Tanjim, Kushal Kafle, Stefan Lee arxiv

Multimodal large language models (MLLMs) are increasingly used to evaluate text-to-image (TTI) generation systems, providing automated judgments based on visual and textual context. However, these "judge" models often su…

Image Clustering

IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments

2025-12-22 · Xu Liu, Yu Liu, Hanshuo Qiu, Yang Qirong 외 arxiv

Vision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robo…

Vision-Language NavigationMultimodal ReasoningData Augmentation

NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities

2025-09-20 · Changyu Zeng, Yifan Wang, Zimu Wang, Wei Wang 외 arxiv

Recent advancements in 2D multimodal large language models (MLLMs) have significantly improved performance in vision-language tasks. However, extending these capabilities to 3D environments remains a distinct challenge d…

Spatial Reasoning

Human Motion Estimation with Everyday Wearables

2025-12-24 · Siqi Zhu, Yixuan Li, Junfu Li, Qi Wu 외 arxiv

While on-body device-based human motion estimation is crucial for applications such as XR interaction, existing methods often suffer from poor wearability, expensive hardware, and cumbersome calibration, which hinder the…

V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D Annotations

2024-12-16 · Jin-Cheng Jhang, Tao Tu, Fu-En Wang, Ke Zhang 외

The field of indoor monocular 3D object detection is gaining significant attention, fueled by the increasing demand in VR/AR and robotic applications. However, its advancement is impeded by the limited availability and d…

3D Object DetectionDepth EstimationMonocular 3D Object DetectionMonocular Depth Estimation+2