Robust Region Feature Synthesizer for Zero-Shot Object Detection
Zero-shot object detection aims at incorporating class semantic vectors to realize the detection of (both seen and) unseen classes given an unconstrained test image. In this study, we reveal the core challenges in this research area: how to synthesize robust region features (for unseen objects) that are as intra-class diverse and inter-class separable as the real samples, so that strong unseen object detectors can be trained upon them. To address these challenges, we build a novel zero-shot object detection framework that contains an Intra-class Semantic Diverging component and an Inter-class Structure Preserving component. The former is used to realize the one-to-more mapping to obtain diverse visual features from each class semantic vector, preventing miss-classifying the real unseen objects as image backgrounds. While the latter is used to avoid the synthesized features too scattered to mix up the inter-class and foreground-background relationship. To demonstrate the effectiveness of the proposed approach, comprehensive experiments on PASCAL VOC, COCO, and DIOR datasets are conducted. Notably, our approach achieves the new state-of-the-art performance on PASCAL VOC and COCO and it is the first study to carry out zero-shot object detection in remote sensing imagery.
Code (1)
Tasks
Generalized Zero-Shot Object DetectionObjectobject-detectionObject DetectionZero-Shot Object DetectionSimilar Papers 제목 키워드 기반
SeeDS: Semantic Separable Diffusion Synthesizer for Zero-shot Food Detection
Food detection is becoming a fundamental task in food computing that supports various multimedia applications, including food recommendation and dietary monitoring. To deal with real-world scenarios, food detection needs…
DenoisingFood recommendationGeneralized Zero-Shot Object DetectionGeneralized Zero-Shot Object Detection on MS-COCO+2GTNet: Generative Transfer Network for Zero-Shot Object Detection
We propose a Generative Transfer Network (GTNet) for zero shot object detection (ZSD). GTNet consists of an Object Detection Module and a Knowledge Transfer Module. The Object Detection Module can learn large-scale seen …
Generative Adversarial NetworkObjectobject-detectionObject Detection+2Transferring neural speech waveform synthesizers to musical instrument sounds generation
Recent neural waveform synthesizers such as WaveNet, WaveGlow, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different methods of waveform generation. The similari…
Audio GenerationAudio SynthesisSpeech SynthesisZero-Shot LearningSynthesizing Knowledge-enhanced Features for Real-world Zero-shot Food Detection
Food computing brings various perspectives to computer vision like vision-based food analysis for nutrition and health. As a fundamental task in food computing, food detection needs Zero-Shot Detection (ZSD) on novel uns…
AttributeGeneralized Zero-Shot Object DetectionNutritionZero-Shot Object DetectionSYNTHONY: A Stress-Aware, Intent-Conditioned Agent for Deep Tabular Generative Models Selection
Deep generative models for tabular data (GANs, diffusion models, and LLM-based generators) exhibit highly non-uniform behavior across datasets; the best-performing synthesizer family depends strongly on distributional st…