When Few-Shot Learning Meets Video Object Detection
Different from static images, videos contain additional temporal and spatial information for better object detection. However, it is costly to obtain a large number of videos with bounding box annotations that are required for supervised deep learning. Although humans can easily learn to recognize new objects by watching only a few video clips, deep learning usually suffers from overfitting. This leads to an important question: how to effectively learn a video object detector from only a few labeled video clips? In this paper, we study the new problem of few-shot learning for video object detection. We first define the few-shot setting and create a new benchmark dataset for few-shot video object detection derived from the widely used ImageNet VID dataset. We employ a transfer-learning framework to effectively train the video object detector on a large number of base-class objects and a few video clips of novel-class objects. By analyzing the results of two methods under this framework (Joint and Freeze) on our designed weak and strong base datasets, we reveal insufficiency and overfitting problems. A simple but effective method, called Thaw, is naturally developed to trade off the two problems and validate our analysis. Extensive experiments on our proposed benchmark datasets with different scenarios demonstrate the effectiveness of our novel analysis in this new few-shot video object detection problem.
Code (0)
등록된 구현이 없습니다.
Tasks
Few-Shot LearningFew-Shot Video Object DetectionObjectobject-detectionObject DetectionTransfer LearningVideo Object DetectionSimilar Papers 제목 키워드 기반
When SAM2 Meets Video Shadow and Mirror Detection
As the successor to the Segment Anything Model (SAM), the Segment Anything Model 2 (SAM2) not only improves performance in image segmentation but also extends its capabilities to video segmentation. However, its effectiv…
Image SegmentationMirror DetectionSegmentationSemantic Segmentation+4When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation
This study investigates the application and performance of the Segment Anything Model 2 (SAM2) in the challenging task of video camouflaged object segmentation (VCOS). VCOS involves detecting objects that blend seamlessl…
Camouflaged Object SegmentationSemantic SegmentationSegment Anything Meets Point Tracking
The Segment Anything Model (SAM) has established itself as a powerful zero-shot image segmentation model, enabled by efficient point-centric annotation and prompt-based models. While click and brush interactions are both…
Interactive Video Object SegmentationObjectPoint TrackingSegmentation+4When SAM Meets Shadow Detection
As a promptable generic object segmentation model, segment anything model (SAM) has recently attracted significant attention, and also demonstrates its powerful performance. Nevertheless, it still meets its Waterloo when…
Image SegmentationMedical Image SegmentationObjectobject-detection+4Single Shot Video Object Detector
Single shot detectors that are potentially faster and simpler than two-stage detectors tend to be more applicable to object detection in videos. Nevertheless, the extension of such object detectors from image to video is…
GPUObjectobject-detectionObject Detection+1