Low-shot Visual Recognition by Shrinking and Hallucinating Features
Low-shot visual learning---the ability to recognize novel object categories from very few examples---is a hallmark of human visual intelligence. Existing machine learning approaches fail to generalize in the same way. To make progress on this foundational problem, we present a low-shot learning benchmark on complex images that mimics challenges faced by recognition systems in the wild. We then propose a) representation regularization techniques, and b) techniques to hallucinate additional training examples for data-starved classes. Together, our methods improve the effectiveness of convolutional networks in low-shot learning, improving the one-shot accuracy on novel classes by 2.3x on the challenging ImageNet dataset.
Code (4)
Tasks
BIG-bench Machine LearningFew-Shot Image ClassificationSimilar Papers 제목 키워드 기반
Temporal Hallucinating for Action Recognition With Few Still Images
Action recognition in still images has been recently promoted by deep learning. However, the success of these deep models heavily depends on huge amount of training images for various action categories, which may not be …
Action RecognitionAction Recognition In Still ImagesDomain AdaptationTemporal Action LocalizationHallucinating Agnostic Images to Generalize Across Domains
The ability to generalize across visual domains is crucial for the robustness of artificial recognition systems. Although many training sources may be available in real contexts, the access to even unlabeled target sampl…
Domain AdaptationDomain GeneralizationUnsupervised Domain AdaptationTracking the Untrackable
Although short-term fully occlusion happens rare in visual object tracking, most trackers will fail under these circumstances. However, humans can still catch up the target by anticipating the trajectory of the target ev…
Object TrackingVisual Object TrackingAdaTranS: Adapting with Boundary-based Shrinking for End-to-End Speech Translation
To alleviate the data scarcity problem in End-to-end speech translation (ST), pre-training on data for speech recognition and machine translation is considered as an important technique. However, the modality gap between…
Machine Translationspeech-recognitionSpeech RecognitionTranslationFew-Shot Learning via Saliency-guided Hallucination of Samples
Learning new concepts from a few of samples is a standard challenge in computer vision. The main directions to improve the learning ability of few-shot training models include (i) a robust similarity learning and (ii) ge…
Few-Shot LearningHallucination