paper-with-me

홈 › Papers

Relating the fundamental frequency of speech with EEG using a dilated convolutional network

2022-07-05 · Corentin Puffay, Jana Van Canneyt, Jonas Vanthornhout, Hugo Van hamme, Tom Francart

To investigate how speech is processed in the brain, we can model the relation between features of a natural speech signal and the corresponding recorded electroencephalogram (EEG). Usually, linear models are used in regression tasks. Either EEG is predicted, or speech is reconstructed, and the correlation between predicted and actual signal is used to measure the brain's decoding ability. However, given the nonlinear nature of the brain, the modeling ability of linear models is limited. Recent studies introduced nonlinear models to relate the speech envelope to EEG. We set out to include other features of speech that are not coded in the envelope, notably the fundamental frequency of the voice (f0). F0 is a higher-frequency feature primarily coded at the brainstem to midbrain level. We present a dilated-convolutional model to provide evidence of neural tracking of the f0. We show that a combination of f0 and the speech envelope improves the performance of a state-of-the-art envelope-based model. This suggests the dilated-convolutional model can extract non-redundant information from both f0 and the envelope. We also show the ability of the dilated-convolutional model to generalize to subjects not included during training. This latter finding will accelerate f0-based hearing diagnosis.

📄 PDF Abstract BibTeX arXiv:2207.01963

Code (0)

등록된 구현이 없습니다.

Tasks

EEGElectroencephalogram (EEG)

Similar Papers 제목 키워드 기반

Deep Interaction between Masking and Mapping Targets for Single-Channel Speech Enhancement

2021-06-09 · Lu Zhang, Mingjiang Wang, Zehua Zhang, Xuyi Zhuang

The most recent deep neural network (DNN) models exhibit impressive denoising performance in the time-frequency (T-F) magnitude domain. However, the phase is also a critical component of the speech signal that is easily …

DenoisingSpeech Enhancement

Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation

2018-09-20 · Yi Luo, Nima Mesgarani

Single-channel, speaker-independent speech separation methods have recently seen great progress. However, the accuracy, latency, and computational cost of such methods remain insufficient. The majority of the previous me…

Multi-task Audio Source SeperationMusic Source SeparationSpeaker SeparationSpeech Enhancement+1

Quasi-Periodic WaveNet: An Autoregressive Raw Waveform Generative Model with Pitch-dependent Dilated Convolution Neural Network

2020-07-11 · Yi-Chiao Wu, Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi 외

In this paper, a pitch-adaptive waveform generative model named Quasi-Periodic WaveNet (QPNet) is proposed to improve the limited pitch controllability of vanilla WaveNet (WN) using pitch-dependent dilated convolution ne…

Tdcgan: Temporal Dilated Convolutional Generative Adversarial Network for End-to-end Speech Enhancement

2020-09-30

In this paper, in order to further deal with the performance degradation caused by ignoring the phase information in conventional speech enhancement systems, we proposed a temporal dilated convolutional generative advers…

Generative Adversarial NetworkSpeech Enhancement

DEEPF0: End-To-End Fundamental Frequency Estimation for Music and Speech Signals

2021-02-11 · Satwinder Singh, Ruili Wang, Yuanhang Qiu

We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. F0 estimation is important in various speech proces…

Information RetrievalMusic Information RetrievalRetrieval