What Is Considered Complete for Visual Recognition?
This is an opinion paper. We hope to deliver a key message that current visual recognition systems are far from complete, i.e., recognizing everything that human can recognize, yet it is very unlikely that the gap can be bridged by continuously increasing human annotations. Based on the observation, we advocate for a new type of pre-training task named learning-by-compression. The computational models (e.g., a deep network) are optimized to represent the visual data using compact features, and the features preserve the ability to recover the original data. Semantic annotations, when available, play the role of weak supervision. An important yet challenging issue is the evaluation of image recovery, where we suggest some design principles and future research directions. We hope our proposal can inspire the community to pursue the compression-recovery tradeoff rather than the accuracy-complexity tradeoff.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Pushing the envelope in deep visual recognition for mobile platforms
Image classification is the task of assigning to an input image a label from a fixed set of categories. One of its most important applicative fields is that of robotics, in particular the needing of a robot to be aware o…
ClassificationGeneral Classificationimage-classificationImage Classification+2On Attention Models for Human Activity Recognition
Most approaches that model time-series data in human activity recognition based on body-worn sensing (HAR) use a fixed size temporal context to represent different activities. This might, however, not be apt for sets of …
Activity RecognitionHuman Activity RecognitionTime SeriesTime Series AnalysisVASR: Visual Analogies of Situation Recognition
A core process in human cognition is analogical mapping: the ability to identify a similar relational structure between different situations. We introduce a novel task, Visual Analogies of Situation Recognition, adapting…
Common Sense ReasoningTripletVisual AnalogiesVisual Commonsense Reasoning+1Redefining Binarization and the Visual Archetype
Although binarization is considered passe, it still remains a highly popular research topic. In this paper we propose a rethinking of what binarization is. We introduce the notion of the visual archetype as the ideal for…
BinarizationHey Human, If your Facial Emotions are Uncertain, You Should Use Bayesian Neural Networks!
Facial emotion recognition is the task to classify human emotions in face images. It is a difficult task due to high aleatoric uncertainty and visual ambiguity. A large part of the literature aims to show progress by inc…
Emotion RecognitionFacial Emotion Recognition