SNIDA: Unlocking Few-Shot Object Detection with Non-linear Semantic Decoupling Augmentation
Once only a few-shot annotated samples are available the performance of learning-based object detection would be heavily dropped. Many few-shot object detection (FSOD) methods have been proposed to tackle this issue by adopting image-level augmentations in linear manners. Nevertheless those handcrafted enhancements often suffer from limited diversity and lack of semantic awareness resulting in unsatisfactory performance. To this end we propose a Semantic-guided Non-linear Instance-level Data Augmentation method (SNIDA) for FSOD by decoupling the foreground and background to increase their diversities respectively. We design a semantic awareness enhancement strategy to separate objects from backgrounds. Concretely masks of instances are extracted by an unsupervised semantic segmentation module. Then the diversity of samples would be improved by fusing instances into different backgrounds. Considering the shortcomings of augmenting images in a limited transformation space of existing traditional data augmentation methods we introduce an object reconstruction enhancement module. The aim of this module is to generate sufficient diversity and non-linear training data at the instance level through a semantic-guided masked autoencoder. In this way the potential of data can be fully exploited in various object detection scenarios. Extensive experiments on PASCAL VOC and MS-COCO demonstrate that the proposed method outperforms baselines by a large margin and achieves new state-of-the-art results under different shot settings.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationDiversityFew-Shot Object DetectionObjectobject-detectionObject DetectionObject ReconstructionSemantic SegmentationUnsupervised Semantic SegmentationSimilar Papers 제목 키워드 기반
Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection
Zero-shot human-object interaction (HOI) detector is capable of generalizing to HOI categories even not encountered during training. Inspired by the impressive zero-shot capabilities offered by CLIP, latest methods striv…
Human-Object Interaction DetectionZero-Shot Human-Object Interaction DetectionMeta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment
Few-shot object detection (FSOD) aims to detect objects using only a few examples. How to adapt state-of-the-art object detectors to the few-shot domain remains challenging. Object proposal is a key ingredient in modern …
Few-Shot LearningFew-Shot Object DetectionMeta-LearningMetric Learning+3Multimodal Reference Visual Grounding
Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding performance by training large models with …
Few-Shot Object DetectionVisual GroundingUnlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
Open-vocabulary 3D object detection (OV-3DDet) aims to localize and recognize both seen and previously unseen object categories within any new 3D scene. While language and vision foundation models have achieved success i…
3D Object DetectionObjectobject-detectionObject DetectionComparison Network for One-Shot Conditional Object Detection
The current advances in object detection depend on large-scale datasets to get good performance. However, there may not always be sufficient samples in many scenarios, which leads to the research on few-shot detection as…
Objectobject-detectionObject Detection