paper-with-me

홈 › Papers

Frequency-Directional Attention Model for Multilingual Automatic Speech Recognition

2022-03-29 · Akihiro Dobashi, Chee Siang Leow, Hiromitsu Nishizaki

This paper proposes a model for transforming speech features using the frequency-directional attention model for End-to-End (E2E) automatic speech recognition. The idea is based on the hypothesis that in the phoneme system of each language, the characteristics of the frequency bands of speech when uttering them are different. By transforming the input Mel filter bank features with an attention model that characterizes the frequency direction, a feature transformation suitable for ASR in each language can be expected. This paper introduces a Transformer-encoder as a frequency-directional attention model. We evaluated the proposed method on a multilingual E2E ASR system for six different languages and found that the proposed method could achieve, on average, 5.3 points higher accuracy than the ASR model for each language by introducing the frequency-directional attention mechanism. Furthermore, visualization of the attention weights based on the proposed method suggested that it is possible to transform acoustic features considering the frequency characteristics of each language.

📄 PDF Abstract BibTeX arXiv:2203.15473

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR

2025-06-16 · Yizhou Peng, Hexin Liu, Eng Siong Chng

This paper introduces the integration of language-specific bi-directional context into a speech large language model (SLLM) to improve multilingual continuous conversational automatic speech recognition (ASR). We propose…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

Amrita@LT-EDI-EACL2021: Hope Speech Detection on Multilingual Text

2021-04-01 · EACL (LTEDI) 2021 4 · Thara S, Ravi teja Tasubilli, Kothamasu Sai rahul

Analysis and deciphering code-mixed data is imperative in academia and industry, in a multilingual country like India, in order to solve problems apropos Natural Language Processing. This paper proposes a bidirectional l…

Hope Speech Detection

Speech Emotion Recognition via an Attentive Time-Frequency Neural Network

2022-10-22 · Cheng Lu, Wenming Zheng, Hailun Lian, Yuan Zong 외

Spectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time-frequency pattern of speech signal for speech emotion recognition (SER). \textcolor{black}{Generally, different e…

Emotion RecognitionSpeech Emotion Recognition

ABARUAH at SemEval-2019 Task 5 : Bi-directional LSTM for Hate Speech Detection

2019-06-01 · SEMEVAL 2019 6 · Arup Baruah, Ferdous Barbhuiya, Kuntal Dey

In this paper, we present the results obtained using bi-directional long short-term memory (BiLSTM) with and without attention and Logistic Regression (LR) models for SemEval-2019 Task 5 titled {''}HatEval: Multilingual …

Hate Speech Detection

Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence Modeling

2021-06-21 · NeurIPS 2021 12 · Hongyu Gong, Yun Tang, Juan Pino, Xian Li

Multi-head attention has each of the attention heads collect salient information from different parts of an input sequence, making it a powerful mechanism for sequence modeling. Multilingual and multi-domain learning are…

speech-recognitionSpeech RecognitionSpeech-to-TextSpeech-to-Text Translation+1