paper-with-me

홈 › Papers

OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D Data

2023-10-20 · Yijie Zhou, Likun Cai, Xianhui Cheng, Zhongxue Gan, xiangyang xue, Wenchao Ding

In the era of big data and large models, automatic annotating functions for multi-modal data are of great significance for real-world AI-driven applications, such as autonomous driving and embodied AI. Unlike traditional closed-set annotation, open-vocabulary annotation is essential to achieve human-level cognition capability. However, there are few open-vocabulary auto-labeling systems for multi-modal 3D data. In this paper, we introduce OpenAnnotate3D, an open-source open-vocabulary auto-labeling system that can automatically generate 2D masks, 3D masks, and 3D bounding box annotations for vision and point cloud data. Our system integrates the chain-of-thought capabilities of Large Language Models (LLMs) and the cross-modality capabilities of vision-language models (VLMs). To the best of our knowledge, OpenAnnotate3D is one of the pioneering works for open-vocabulary multi-modal 3D auto-labeling. We conduct comprehensive evaluations on both public and in-house real-world datasets, which demonstrate that the system significantly improves annotation efficiency compared to manual annotation while providing accurate open-vocabulary auto-annotating results.

📄 PDF Abstract BibTeX arXiv:2310.13398

Code (1)

fudan-projecttitan/openannotate3d 공식 구현 pytorch

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

OVI-MAP:Open-Vocabulary Instance-Semantic Mapping

2026-03-27 · Zilong Deng, Federico Tombari, Marc Pollefeys, Johanna Wald 외 arxiv

Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the need for robust instance segmentation, re…

Instance Segmentation

Region-centric Image-Language Pretraining for Open-Vocabulary Detection

2023-09-29 · Dahun Kim, Anelia Angelova, Weicheng Kuo

We present a new open-vocabulary detection approach based on region-centric image-language pretraining to bridge the gap between image-level pretraining and open-vocabulary object detection. At the pretraining phase, we …

Contrastive LearningObjectobject-detectionObject Detection+2

VESPA: Towards un(Human)supervised Open-World Pointcloud Labeling for Autonomous Driving

2025-07-27 · Levente Tempfli, Esteban Rivera, Markus Lienkamp arxiv

Data collection for autonomous driving is rapidly accelerating, but manual annotation, especially for 3D labels, remains a major bottleneck due to its high cost and labor intensity. Autolabeling has emerged as a scalable…

Scene UnderstandingAutonomous DrivingObject DetectionPoint Clouds

Learning Pseudo-Labeler beyond Noun Concepts for Open-Vocabulary Object Detection

2023-12-04 · Sunghun Kang, Junbum Cha, Jonghwan Mun, Byungseok Roh 외

Open-vocabulary object detection (OVOD) has recently gained significant attention as a crucial step toward achieving human-like visual intelligence. Existing OVOD methods extend target vocabulary from pre-defined categor…

Image to textobject-detectionObject DetectionOpen-vocabulary object detection+3

A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and Future

2023-07-18 · Chaoyang Zhu, Long Chen

As the most fundamental scene understanding tasks, object detection and segmentation have made tremendous progress in deep learning era. Due to the expensive manual labeling cost, the annotated categories in existing dat…

Knowledge Distillationobject-detectionObject DetectionPanoptic Segmentation+4