paper-with-me

홈 › Papers

QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection

2021-03-16 · CVPR 2022 1 · Chenhongyi Yang, Zehao Huang, Naiyan Wang

While general object detection with deep learning has achieved great success in the past few years, the performance and efficiency of detecting small objects are far from satisfactory. The most common and effective way to promote small object detection is to use high-resolution images or feature maps. However, both approaches induce costly computation since the computational cost grows squarely as the size of images and features increases. To get the best of two worlds, we propose QueryDet that uses a novel query mechanism to accelerate the inference speed of feature-pyramid based object detectors. The pipeline composes two steps: it first predicts the coarse locations of small objects on low-resolution features and then computes the accurate detection results using high-resolution features sparsely guided by those coarse positions. In this way, we can not only harvest the benefit of high-resolution feature maps but also avoid useless computation for the background area. On the popular COCO dataset, the proposed method improves the detection mAP by 1.0 and mAP-small by 2.0, and the high-resolution inference speed is improved to 3.0x on average. On VisDrone dataset, which contains more small objects, we create a new state-of-the-art while gaining a 2.3x high-resolution acceleration on average. Code is available at https://github.com/ChenhongyiYang/QueryDet-PyTorch.

📄 PDF Abstract BibTeX arXiv:2103.09136

Code (1)

ChenhongyiYang/QueryDet-PyTorch 공식 구현 pytorch

Tasks

object-detectionObject DetectionSmall Object DetectionVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference

2025-10-20 · Samir Khaki, Junxian Guo, Jiaming Tang, Shang Yang 외 arxiv

Vision Language Models (VLMs) have rapidly advanced in integrating visual and textual reasoning, powering applications across high-resolution image understanding, long-video analysis, and multi-turn conversation. However…

QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill

2026-02-09 · Dalton Jones, Junyoung Park, Matthew Morse, Mingu Lee 외 arxiv

We present QUOKA: Query-oriented KV selection for efficient attention, a training-free and hardware agnostic sparse attention algorithm for accelerating transformer inference under chunked prefill. While many queries foc…

H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing

2025-06-20 · Akanksha Singh, Yi-Ping Phoebe Chen, Vipul Arora

Query-by-example spoken term detection (QbE-STD) searches for matching words or phrases in an audio dataset using a sample spoken query. When annotated data is limited or unavailable, QbE-STD is often done using template…

Dynamic Time WarpingRepresentation LearningRetrievalTemplate Matching

IoU-Enhanced Attention for End-to-End Task Specific Object Detection

2022-09-21 · Jing Zhao, Shengjian Wu, Li Sun, Qingli Li

Without densely tiled anchor boxes or grid points in the image, sparse R-CNN achieves promising results through a set of object queries and proposal boxes updated in the cascaded training manner. However, due to the spar…

Objectobject-detectionObject Detection

Fourier-Mixed Window Attention: Accelerating Informer for Long Sequence Time-Series Forecasting

2023-07-02 · Nhat Thanh Tran, Jack Xin

We study a fast local-global window-based attention method to accelerate Informer for long sequence time-series forecasting. While window attention being local is a considerable computational saving, it lacks the ability…

Time SeriesTime Series Forecasting