paper-with-me

홈 › Papers

Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

2024-05-16 · Tianhe Ren, Qing Jiang, Shilong Liu, Zhaoyang Zeng, Wenlong Liu, Han Gao, Hongjie Huang, Zhengyu Ma, Xiaoke Jiang, Yihao Chen, Yuda Xiong, Hao Zhang, Feng Li, Peijun Tang, Kent Yu, Lei Zhang

This paper introduces Grounding DINO 1.5, a suite of advanced open-set object detection models developed by IDEA Research, which aims to advance the "Edge" of open-set object detection. The suite encompasses two models: Grounding DINO 1.5 Pro, a high-performance model designed for stronger generalization capability across a wide range of scenarios, and Grounding DINO 1.5 Edge, an efficient model optimized for faster speed demanded in many applications requiring edge deployment. The Grounding DINO 1.5 Pro model advances its predecessor by scaling up the model architecture, integrating an enhanced vision backbone, and expanding the training dataset to over 20 million images with grounding annotations, thereby achieving a richer semantic understanding. The Grounding DINO 1.5 Edge model, while designed for efficiency with reduced feature scales, maintains robust detection capabilities by being trained on the same comprehensive dataset. Empirical results demonstrate the effectiveness of Grounding DINO 1.5, with the Grounding DINO 1.5 Pro model attaining a 54.3 AP on the COCO detection benchmark and a 55.7 AP on the LVIS-minival zero-shot transfer benchmark, setting new records for open-set object detection. Furthermore, the Grounding DINO 1.5 Edge model, when optimized with TensorRT, achieves a speed of 75.2 FPS while attaining a zero-shot performance of 36.2 AP on the LVIS-minival benchmark, making it more suitable for edge computing scenarios. Model examples and demos with API will be released at https://github.com/IDEA-Research/Grounding-DINO-1.5-API

📄 PDF Abstract BibTeX arXiv:2405.10300

Code (3)

idea-research/grounding-dino-1.5-api 공식 구현
idea-research/grounded-sam-2 pytorch
mit-han-lab/efficientvit pytorch

Tasks

Edge-computingFew-Shot Object Detectionobject-detectionObject DetectionZero-Shot Object Detection

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Multi-Head Attention 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

2023-03-09 · Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li 외

In this paper, we present an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category …

DecoderObject DetectionReferring ExpressionReferring Expression Comprehension+2

DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

2024-11-21 · Tianhe Ren, Yihao Chen, Qing Jiang, Zhaoyang Zeng 외

In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X employs the same Transformer-based encod…

Long-tailed Object DetectionObjectobject-detectionObject Detection+3

An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

2024-01-04 · Xiangyu Zhao, Yicheng Chen, Shilin Xu, Xiangtai Li 외

Grounding-DINO is a state-of-the-art open-set detection model that tackles multiple vision tasks including Open-Vocabulary Detection (OVD), Phrase Grounding (PG), and Referring Expression Comprehension (REC). Its effecti…

Described Object DetectionPhrase GroundingReferring ExpressionReferring Expression Comprehension

Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection

2025-07-23 · Yehao Lu, Minghe Weng, Zekang Xiao, Rui Jiang 외 arxiv

The Mixture of Experts (MoE) architecture has excelled in Large Vision-Language Models (LVLMs), yet its potential in real-time open-vocabulary object detectors, which also leverage large-scale vision-language datasets bu…

Object Detection

HDINO: A Concise and Efficient Open-Vocabulary Detector

2026-03-03 · Hao Zhang, Yiqun Wang, Qinran Lin, Runze Fan 외 arxiv

Despite the growing interest in open-vocabulary object detection in recent years, most existing methods rely heavily on manually curated fine-grained training datasets as well as resource-intensive layer-wise cross-modal…

Object Detection