A Comparative Attention Framework for Better Few-Shot Object Detection on Aerial Images
Few-Shot Object Detection (FSOD) methods are mainly designed and evaluated on natural image datasets such as Pascal VOC and MS COCO. However, it is not clear whether the best methods for natural images are also the best for aerial images. Furthermore, direct comparison of performance between FSOD methods is difficult due to the wide variety of detection frameworks and training strategies. Therefore, we propose a benchmarking framework that provides a flexible environment to implement and compare attention-based FSOD methods. The proposed framework focuses on attention mechanisms and is divided into three modules: spatial alignment, global attention, and fusion layer. To remain competitive with existing methods, which often leverage complex training, we propose new augmentation techniques designed for object detection. Using this framework, several FSOD methods are reimplemented and compared. This comparison highlights two distinct performance regimes on aerial and natural images: FSOD performs worse on aerial images. Our experiments suggest that small objects, which are harder to detect in the few-shot setting, account for the poor performance. Finally, we develop a novel multiscale alignment method, Cross-Scales Query-Support Alignment (XQSA) for FSOD, to improve the detection of small objects. XQSA outperforms the state-of-the-art significantly on DOTA and DIOR.
Code (1)
Tasks
BenchmarkingFew-Shot Object Detectionobject-detectionObject DetectionSimilar Papers 제목 키워드 기반
GATE-AD: Graph Attention Network Encoding For Few-Shot Industrial Visual Anomaly Detection
Few-Shot Industrial Visual Anomaly Detection (FS-IVAD) comprises a critical task in modern manufacturing settings, where automated product inspection systems need to identify rare defects using only a handful of normal/d…
Anomaly DetectionAttentional Prototype Inference for Few-Shot Segmentation
This paper aims to address few-shot segmentation. While existing prototype-based methods have achieved considerable success, they suffer from uncertainty and ambiguity caused by limited labeled examples. In this work, we…
Bayesian InferenceFew-Shot Semantic SegmentationSegmentationSemantic SegmentationNeural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping
One-shot transfer of dexterous grasps to novel scenes with object and context variations has been a challenging problem. While distilled feature fields from large vision models have enabled semantic correspondences acros…
DecoderObject Counting with GPT-4o and GPT-5: A Comparative Study
Zero-shot object counting attempts to estimate the number of object instances belonging to novel categories that the vision model performing the counting has never encountered during training. Existing methods typically …
Object CountingAttention Based Simple Primitives for Open World Compositional Zero-Shot Learning
Compositional Zero-Shot Learning (CZSL) aims to predict unknown compositions made up of attribute and object pairs. Predicting compositions unseen during training is a challenging task. We are exploring Open World Compos…
AttributeCompositional Zero-Shot LearningObjectZero-Shot Learning