On the Exploration of Convolutional Fusion Networks for Visual Recognition
Despite recent advances in multi-scale deep representations, their limitations are attributed to expensive parameters and weak fusion modules. Hence, we propose an efficient approach to fuse multi-scale deep representations, called convolutional fusion networks (CFN). Owing to using 1$\times$1 convolution and global average pooling, CFN can efficiently generate the side branches while adding few parameters. In addition, we present a locally-connected fusion module, which can learn adaptive weights for the side branches and form a discriminatively fused feature. CFN models trained on the CIFAR and ImageNet datasets demonstrate remarkable improvements over the plain CNNs. Furthermore, we generalize CFN to three new tasks, including scene recognition, fine-grained recognition and image retrieval. Our experiments show that it can obtain consistent improvements towards the transferring tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
Image RetrievalRetrievalScene RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Exploring Deep Learning for Joint Audio-Visual Lip Biometrics
Audio-visual (AV) lip biometrics is a promising authentication technique that leverages the benefits of both the audio and visual modalities in speech communication. Previous works have demonstrated the usefulness of AV …
Deep LearningSpeaker RecognitionMaking Convolutional Networks Recurrent for Visual Sequence Learning
Recurrent neural networks (RNNs) have emerged as a powerful model for a broad range of machine learning problems that involve sequential data. While an abundance of work exists to understand and improve RNNs in the conte…
Action RecognitionFace AlignmentGesture RecognitionHand Gesture Recognition+6Cloud based Scalable Object Recognition from Video Streams using Orientation Fusion and Convolutional Neural Networks
Object recognition from live video streams comes with numerous challenges such as the variation in illumination conditions and poses. Convolutional neural networks (CNNs) have been widely used to perform intelligent visu…
ObjectObject RecognitionLip Graph Assisted Audio-Visual Speech Recognition Using Bidirectional Synchronous Fusion
Current studies have shown that extracting representative visual features and efficiently fusing audio and visual modalities are vital for audio-visual speech recognition (AVSR), but these are still challenging. To this …
Audio-Visual Speech RecognitionLandmark-based Lipreadingspeech-recognitionSpeech Recognition+1Combining Multiple Views for Visual Speech Recognition
Visual speech recognition is a challenging research problem with a particular practical application of aiding audio speech recognition in noisy scenarios. Multiple camera setups can be beneficial for the visual speech re…
Sentencespeech-recognitionSpeech RecognitionVisual Speech Recognition