paper-with-me

홈 › Papers

LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection

2026-06-24 · Qinzhe Yang, Dongyu Wang, Haohan Niu, Jia Xu, Zhenwei Shi, Zhengxia Zou arxiv

Remote sensing object detection has advanced rapidly with the development of large-scale benchmarks and modern detection architectures. However, existing datasets and detectors remain fragmented. Most benchmarks focus on limited categories, fixed spatial resolutions, or a single sensor, while detectors still struggle to work across different sensors and categorical systems. In this paper, we introduce LEVIRDet-159, the largest and most comprehensive remote sensing object detection dataset to date, with 159 categories, 2.56 million bounding boxes, and 700k fine-grained annotations under a multi-level taxonomy. In each key scale dimension, LEVIRDet-159 exceeds the corresponding largest existing remote sensing object detection dataset, containing approximately (7x) more images, (6x) more object instances, and (4x) more categories. Based on this dataset, we design LEVIRDetNet, a scale-hierarchy-aware detection foundation model for universal remote sensing object detection. LEVIRDetNet couples online visual Ground Sampling Distance (GSD) prediction, GSD-conditioned query modulation and allocation, and a hierarchy-aware detection head for mixed-granularity remote sensing supervision. Under stringent evaluation settings, LEVIRDetNet demonstrates strong cross-domain generalization. Even without target-domain training or fine-tuning, it achieves state-of-the-art detection performance on 9 external benchmarks, improving the strongest fully supervised competing methods by 5.02 mAP on average under each benchmark's primary metric. We hope this study will facilitate the development of strongly generalizable remote sensing object detection across diverse category systems, spatial resolutions, and sensor platforms. The dataset and trained models will be released at https://qinzheyang.github.io/LEVIRDet/, accompanying the final paper.

📄 PDF Abstract BibTeX arXiv:2606.25312

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationObject Detection

Similar Papers 제목 키워드 기반

Collaborative Feature Learning from Social Media

2015-02-05 · CVPR 2015 6 · Chen Fang, Hailin Jin, Jianchao Yang, Zhe Lin

Image feature representation plays an essential role in image recognition and related tasks. The current state-of-the-art feature learning paradigm is supervised learning from labeled data. However, this paradigm require…

Large-Scale Category Structure Aware Image Categorization

2011-12-01 · NeurIPS 2011 12 · Bin Zhao, Fei Li, Eric P. Xing

Most previous research on image categorization has focused on medium-scale data sets, while large-scale image categorization with millions of images from thousands of categories remains a challenge. With the emergence of…

Image Categorization

Species196: A One-Million Semi-supervised Dataset for Fine-grained Species Recognition

2023-09-26 · NeurIPS 2023 11

The development of foundation vision models has pushed the general visual recognition to a high level, but cannot well address the fine-grained recognition in specialized domain such as invasive species classification. I…

TEDDY: A Family Of Foundation Models For Understanding Single Cell Biology

2025-03-05 · Alexis Chevalier, Soumya Ghosh, Urvi Awasthi, James Watkins 외

Understanding the biological mechanism of disease is critical for medicine, and in particular drug discovery. AI-powered analysis of genome-scale biological data hold great potential in this regard. The increasing availa…

Drug Discovery

Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks

2023-06-07 · Haiyang Xu, Qinghao Ye, Xuan Wu, Ming Yan 외

To promote the development of Vision-Language Pre-training (VLP) and multimodal Large Language Model (LLM) in the Chinese community, we firstly release the largest public Chinese high-quality video-language dataset named…

Cross-Modal RetrievalLanguage ModellingLarge Language ModelMultimodal Large Language Model+2