The Treasure beneath Convolutional Layers: Cross-convolutional-layer Pooling for Image Classification
A number of recent studies have shown that a Deep Convolutional Neural Network (DCNN) pretrained on a large dataset can be adopted as a universal image description which leads to astounding performance in many visual classification tasks. Most of these studies, if not all, adopt activations of the fully-connected layer of a DCNN as the image or region representation and it is believed that convolutional layer activations are less discriminative. This paper, however, advocates that if used appropriately convolutional layer activations can be turned into a powerful image representation which enjoys many advantages over fully-connected layer activations. This is achieved by adopting a new technique proposed in this paper called cross-convolutional-layer pooling. More specifically, it extracts subarrays of feature maps of one convolutional layer as local features and pools the extracted features with the guidance of feature maps of the successive convolutional layer. Compared with exising methods that apply DCNNs in the local feature setting, the proposed method is significantly faster since it requires much fewer times of DCNN forward computation. Moreover, it avoids the domain mismatch issue which is usually encountered when applying fully connected layer activations to describe local regions. By applying our method to four popular visual classification tasks, it is demonstrated that the proposed method can achieve comparable or in some cases significantly better performance than existing fully-connected layer based image representations while incurring much lower computational cost.
Code (1)
Tasks
General Classificationimage-classificationImage ClassificationImage DescriptionSimilar Papers 제목 키워드 기반
Deep Descriptor Transforming for Image Co-Localization
Reusable model design becomes desirable with the rapid expansion of machine learning applications. In this paper, we focus on the reusability of pre-trained deep convolutional models. Specifically, different from treatin…
Unsupervised Object Discovery and Co-Localization by Deep Descriptor Transforming
Reusable model design becomes desirable with the rapid expansion of computer vision and machine learning applications. In this paper, we focus on the reusability of pre-trained deep convolutional models. Specifically, di…
Objectobject-detectionObject DetectionObject Discovery+2Taxon and trait recognition from digitized herbarium specimens using deep convolutional neural networks
Herbaria worldwide are housing a treasure of 100s of millions of herbarium specimens, which are increasingly being digitized in recent years and thereby made more easily accessible to the scientific community. At the sam…
ManagementThe Treasure Beneath Multiple Annotations: An Uncertainty-aware Edge Detector
Deep learning-based edge detectors heavily rely on pixel-wise labels which are often provided by multiple annotators. Existing methods fuse multiple annotations using a simple voting process, ignoring the inherent ambigu…
DecoderEdge DetectionSelf-Attention Generative Adversarial Network for Speech Enhancement
Existing generative adversarial networks (GANs) for speech enhancement solely rely on the convolution operation, which may obscure temporal dependencies across the sequence input. To remedy this issue, we propose a self-…
Generative Adversarial NetworkSpeech Enhancement