paper-with-me

홈 › Papers

OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object Tracking

2024-10-23 · Haiji Liang, Ruize Han

Open-vocabulary object perception has become an important topic in artificial intelligence, which aims to identify objects with novel classes that have not been seen during training. Under this setting, open-vocabulary object detection (OVD) in a single image has been studied in many literature. However, open-vocabulary object tracking (OVT) from a video has been studied less, and one reason is the shortage of benchmarks. In this work, we have built a new large-scale benchmark for open-vocabulary multi-object tracking namely OVT-B. OVT-B contains 1,048 categories of objects and 1,973 videos with 637,608 bounding box annotations, which is much larger than the sole open-vocabulary tracking dataset, i.e., OVTAO-val dataset (200+ categories, 900+ videos). The proposed OVT-B can be used as a new benchmark to pave the way for OVT research. We also develop a simple yet effective baseline method for OVT. It integrates the motion features for object tracking, which is an important feature for MOT but is ignored in previous OVT methods. Experimental results have verified the usefulness of the proposed benchmark and the effectiveness of our method. We have released the benchmark to the public at https://github.com/Coo1Sea/OVT-B-Dataset.

📄 PDF Abstract BibTeX arXiv:2410.17534

Code (1)

coo1sea/ovt-b-dataset 공식 구현 jax

Tasks

Multi-Object TrackingObjectobject-detectionObject DetectionObject TrackingOpen-vocabulary object detectionOpen Vocabulary Object Detection

Similar Papers 제목 키워드 기반

OVTrack: Open-Vocabulary Multiple Object Tracking

2023-04-17 · CVPR 2023 1 · Siyuan Li, Tobias Fischer, Lei Ke, Henghui Ding 외

The ability to recognize, localize and track dynamic objects in a scene is fundamental to many real-world applications, such as self-driving and robotic systems. Yet, traditional multiple object tracking (MOT) benchmarks…

DenoisingHallucinationKnowledge DistillationMulti-Object Tracking+3

Search3D: Hierarchical Open-Vocabulary 3D Segmentation

2024-09-27 · Ayca Takmaz, Alexandros Delitzas, Robert W. Sumner, Francis Engelmann 외

Open-vocabulary 3D segmentation enables exploration of 3D spaces using free-form text descriptions. Existing methods for open-vocabulary 3D instance segmentation primarily focus on identifying object-level instances but …

3D Instance Segmentation3D Part SegmentationInstance SegmentationObject+3

Object2Scene: Putting Objects in Context for Open-Vocabulary 3D Detection

2023-09-18 · Chenming Zhu, Wenwei Zhang, Tai Wang, Xihui Liu 외

Point cloud-based open-vocabulary 3D object detection aims to detect 3D categories that do not have ground-truth annotations in the training set. It is extremely challenging because of the limited data and annotations (b…

3D Object Detection3D Open-Vocabulary Object DetectionContrastive LearningObject+3

LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors

2024-02-07 · Sheng Jin, Xueying Jiang, Jiaxing Huang, Lewei Lu 외

Inspired by the outstanding zero-shot capability of vision language models (VLMs) in image classification tasks, open-vocabulary object detection has attracted increasing interest by distilling the broad VLM knowledge in…

image-classificationImage Classificationobject-detectionObject Detection+2

OpenUrban3D: Annotation-Free Open-Vocabulary Semantic Segmentation of Large-Scale Urban Point Clouds

2025-09-13 · Chongyu Wang, Kunlei Jing, Jihua Zhu, Di Wang arxiv

Open-vocabulary semantic segmentation enables models to recognize and segment objects from arbitrary natural language descriptions, offering the flexibility to handle novel, fine-grained, or functionally defined categori…

Point Cloud SegmentationSemantic SegmentationScene UnderstandingPoint Clouds