paper-with-me

홈 › Papers

AMFFCN: Attentional Multi-layer Feature Fusion Convolution Network for Audio-visual Speech Enhancement

2021-01-15 · Xinmeng Xu, Jianjun Hao

Audio-visual speech enhancement system is regarded to be one of promising solutions for isolating and enhancing speech of desired speaker. Conventional methods focus on predicting clean speech spectrum via a naive convolution neural network based encoder-decoder architecture, and these methods a) not adequate to use data fully and effectively, b) cannot process features selectively. The proposed model addresses these drawbacks, by a) applying a model that fuses audio and visual features layer by layer in encoding phase, and that feeds fused audio-visual features to each corresponding decoder layer, and more importantly, b) introducing soft threshold attention into the model to select the informative modality softly. This paper proposes attentional audio-visual multi-layer feature fusion model, in which soft threshold attention unit are applied on feature mapping at every layer of decoder. The proposed model demonstrates the superior performance of the network against the state-of-the-art models.

📄 PDF Abstract BibTeX arXiv:2101.06268

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSpeech Enhancement

Similar Papers 제목 키워드 기반

Attentional Feature Fusion

2020-09-29 · Yimian Dai, Fabian Gieseke, Stefan Oehmcke, Yiquan Wu 외

Feature fusion, the combination of features from different layers or branches, is an omnipresent part of modern network architectures. It is often implemented via simple operations, such as summation or concatenation, bu…

Image Classification

Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition

2021-11-09 · R. Gnana Praveen, Eric Granger, Patrick Cardinal

Multimodal analysis has recently drawn much interest in affective computing, since it can improve the overall accuracy of emotion recognition over isolated uni-modal approaches. The most effective techniques for multimod…

Emotion RecognitionMultimodal Emotion Recognition

Deep Multi-Model Fusion for Single-Image Dehazing

2019-10-01 · ICCV 2019 10 · Zijun Deng, Lei Zhu, Xiaowei Hu, Chi-Wing Fu 외

This paper presents a deep multi-model fusion network to attentively integrate multiple models to separate layers and boost the performance in single-image dehazing. To do so, we first formulate the attentional feature i…

Image DehazingmodelSingle Image Dehazing

Bidirectional Multiscale Feature Aggregation for Speaker Verification

2021-04-01 · Jiajun Qi, Wu Guo, Bin Gu

In this paper, we propose a novel bidirectional multiscale feature aggregation (BMFA) network with attentional fusion modules for text-independent speaker verification. The feature maps from different stages of the backb…

Speaker VerificationText-Independent Speaker Verification

Chinese Herbal Recognition based on Competitive Attentional Fusion of Multi-hierarchies Pyramid Features

2018-12-23 · Yingxue Xu, Guihua Wen, Yang Hu, Mingnan Luo 외

Convolution neural netwotks (CNNs) are successfully applied in image recognition task. In this study, we explore the approach of automatic herbal recognition with CNNs and build the standard Chinese herbs datasets firstl…