paper-with-me

홈 › Papers

Attentive Fusion Enhanced Audio-Visual Encoding for Transformer Based Robust Speech Recognition

2020-08-06 · Liangfa Wei, Jie Zhang, JunFeng Hou, Li-Rong Dai

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual alignment and modality reliability. Different from the previous end-to-end approaches where the audio-visual fusion is performed after encoding each modality, in this paper we propose to integrate an attentive fusion block into the encoding process. It is shown that the proposed audio-visual fusion method in the encoder module can enrich audio-visual representations, as the relevance between the two modalities is leveraged. In line with the transformer-based architecture, we implement the embedded fusion block using a multi-head attention based audiovisual fusion with one-way or two-way interactions. The proposed method can sufficiently combine the two streams and weaken the over-reliance on the audio modality. Experiments on the LRS3-TED dataset demonstrate that the proposed method can increase the recognition rate by 0.55%, 4.51% and 4.61% on average under the clean, seen and unseen noise conditions, respectively, compared to the state-of-the-art approach.

📄 PDF Abstract BibTeX arXiv:2008.02686

Code (0)

등록된 구현이 없습니다.

Tasks

Robust Speech Recognitionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Learning Audio-Visual embedding for Person Verification in the Wild

2022-09-09 · Peiwen Sun, Shanshan Zhang, Zishan Liu, Yougen Yuan 외

It has already been observed that audio-visual embedding is more robust than uni-modality embedding for person verification. Here, we proposed a novel audio-visual strategy that considers aggregators from a fusion perspe…

Face Verification

Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition

2024-05-26 · Tong Shi, Xuri Ge, Joemon M. Jose, Nicolas Pugeault 외

Capturing complex temporal relationships between video and audio modalities is vital for Audio-Visual Emotion Recognition (AVER). However, existing methods lack attention to local details, such as facial state changes be…

Emotion RecognitionOptical Flow Estimation

Attentive Filtering Networks for Audio Replay Attack Detection

2018-10-31 · Cheng-I Lai, Alberto Abad, Korin Richmond, Junichi Yamagishi 외

An attacker may use a variety of techniques to fool an automatic speaker verification system into accepting them as a genuine user. Anti-spoofing methods meanwhile aim to make the system robust against such attacks. The …

Speaker Verification

Continuous Emotion Recognition with Audio-visual Leader-follower Attentive Fusion

2021-07-02 · Su Zhang, Yi Ding, Ziquan Wei, Cuntai Guan

We propose an audio-visual spatial-temporal deep neural network with: (1) a visual block containing a pretrained 2D-CNN followed by a temporal convolutional network (TCN); (2) an aural block containing several parallel T…

Emotion Recognition

Diffusion Models as Masked Audio-Video Learners

2023-10-05 · Elvis Nunez, Yanzi Jin, Mohammad Rastegari, Sachin Mehta 외

Over the past several years, the synchronization between audio and visual signals has been leveraged to learn richer audio-visual representations. Aided by the large availability of unlabeled videos, many unsupervised tr…

Audio ClassificationContrastive Learning