paper-with-me

Papers

Cross-View Open-Vocabulary Object Detection in Aerial Imagery

2025-10-04 · Jyoti Kini, Rohit Gupta, Mubarak Shah arxiv

Traditional object detection models are typically trained on a fixed set of classes, limiting their flexibility and making it costly to incorporate new categories. Open-vocabulary object detection addresses this limitation by enabling models to identify unseen classes without explicit training. Leveraging pretrained models contrastively trained on abundantly available ground-view image-text classification pairs provides a strong foundation for open-vocabulary object detection in aerial imagery. Domain shifts, viewpoint variations, and extreme scale differences make direct knowledge transfer across domains ineffective, requiring specialized adaptation strategies. In this paper, we propose a novel framework for adapting open-vocabulary representations from ground-view images to solve object detection in aerial imagery through structured domain alignment. The method introduces contrastive image-to-image alignment to enhance the similarity between aerial and ground-view embeddings and employs multi-instance vocabulary associations to align aerial images with text embeddings. Extensive experiments on the xView, DOTAv2, VisDrone, DIOR, and HRRSD datasets are used to validate our approach. Our open-vocabulary model achieves improvements of +6.32 mAP on DOTAv2, +4.16 mAP on VisDrone (Images), and +3.46 mAP on HRRSD in the zero-shot setting when compared to finetuned closed-vocabulary dataset-specific model performance, thus paving the way for more flexible and scalable object detection systems in aerial applications.

📄 PDF Abstract BibTeX arXiv:2510.03858

Code (0)

등록된 구현이 없습니다.

Tasks

Text ClassificationObject Detection

Similar Papers 제목 키워드 기반

Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection

2026-03-23 · Youbin Kim, Jinho Park, Hogun Park, Eunbyung Park arxiv

Open-vocabulary 3D object detection aims to localize and recognize objects beyond a fixed training taxonomy. In multi-view RGB settings, recent approaches often decouple geometry-based instance construction from semantic…

3D Object Detection

F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

2022-09-30 · Weicheng Kuo, Yin Cui, Xiuye Gu, AJ Piergiovanni 외

We present F-VLM, a simple open-vocabulary object detection method built upon Frozen Vision and Language Models. F-VLM simplifies the current multi-stage training pipeline by eliminating the need for knowledge distillati…

Knowledge Distillationobject-detectionObject DetectionOpen-vocabulary object detection+1

Zoo3D: Zero-Shot 3D Object Detection at Scene Level

2025-11-25 · Andrey Lemeshko, Bulat Gabdullin, Nikita Drozdov, Anton Konushin 외 arxiv

3D object detection is fundamental for spatial understanding. Real-world environments demand models capable of recognizing diverse, previously unseen objects, which remains a major limitation of closed-set methods. Exist…

3D Object DetectionGraph ClusteringPoint Clouds

Sparse Multiview Open-Vocabulary 3D Detection

2025-09-19 · Olivier Moliner, Viktor Larsson, Kalle Åström arxiv

The ability to interpret and comprehend a 3D scene is essential for many vision and robotics systems. In numerous applications, this involves 3D object detection, i.e.~identifying the location and dimensions of objects b…

3D Object Detection

FM-OV3D: Foundation Model-based Cross-modal Knowledge Blending for Open-Vocabulary 3D Detection

2023-12-22 · Dongmei Zhang, Chang Li, Ray Zhang, Shenghao Xie 외

The superior performances of pre-trained foundation models in various visual tasks underscore their potential to enhance the 2D models' open-vocabulary ability. Existing methods explore analogous applications in the 3D s…

3D Object Detection3D Open-Vocabulary Object Detectionobject-detectionObject Detection+1