paper-with-me

홈 › Papers

Analysis of EEG frequency bands for Envisioned Speech Recognition

2022-03-29 · Ayush Tripathi

The use of Automatic speech recognition (ASR) interfaces have become increasingly popular in daily life for use in interaction and control of electronic devices. The interfaces currently being used are not feasible for a variety of users such as those suffering from a speech disorder, locked-in syndrome, paralysis or people with utmost privacy requirements. In such cases, an interface that can identify envisioned speech using electroencephalogram (EEG) signals can be of great benefit. Various works targeting this problem have been done in the past. However, there has been limited work in identifying the frequency bands ($\delta, \theta, \alpha, \beta, \gamma$) of the EEG signal that contribute towards envisioned speech recognition. Therefore, in this work, we aim to analyze the significance of different EEG frequency bands and signals obtained from different lobes of the brain and their contribution towards recognizing envisioned speech. Signals obtained from different lobes and bandpass filtered for different frequency bands are fed to a spatio-temporal deep learning architecture with Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM). The performance is evaluated on a publicly available dataset comprising of three classification tasks - digit, character and images. We obtain a classification accuracy of $85.93\%$, $87.27\%$ and $87.51\%$ for the three tasks respectively. The code for the implementation has been made available at https://github.com/ayushayt/ImaginedSpeechRecognition.

📄 PDF Abstract BibTeX arXiv:2203.15250

Code (1)

ayushayt/imaginedspeechrecognition 공식 구현

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)EEGElectroencephalogram (EEG)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Speech Emotion Recognition via an Attentive Time-Frequency Neural Network

2022-10-22 · Cheng Lu, Wenming Zheng, Hailun Lian, Yuan Zong 외

Spectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time-frequency pattern of speech signal for speech emotion recognition (SER). \textcolor{black}{Generally, different e…

Emotion RecognitionSpeech Emotion Recognition

Frequency-centroid features for word recognition of non-native English speakers

2022-06-14 · Pierre Berjon, Rajib Sharma, Avishek Nag, Soumyabrata Dev

The objective of this work is to investigate complementary features which can aid the quintessential Mel frequency cepstral coefficients (MFCCs) in the task of closed, limited set word recognition for non-native English …

TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech Dereverberation

2021-03-31 · Helin Wang, Bo Wu, LianWu Chen, Meng Yu 외

In this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose a temporal-contextual attention approach…

Room Impulse Response (RIR)Speech Dereverberation

Frequency-Directional Attention Model for Multilingual Automatic Speech Recognition

2022-03-29 · Akihiro Dobashi, Chee Siang Leow, Hiromitsu Nishizaki

This paper proposes a model for transforming speech features using the frequency-directional attention model for End-to-End (E2E) automatic speech recognition. The idea is based on the hypothesis that in the phoneme syst…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Radically Old Way of Computing Spectra: Applications in End-to-End ASR

2021-03-25 · Samik Sadhu, Hynek Hermansky

We propose a technique to compute spectrograms using Frequency Domain Linear Prediction (FDLP) that uses all-pole models to fit the squared Hilbert envelope of speech in different frequency sub-bands. The spectrogram of …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition