Cognitive Coding of Speech
We propose an approach for cognitive coding of speech by unsupervised extraction of contextual representations in two hierarchical levels of abstraction. Speech attributes such as phoneme identity that last one hundred milliseconds or less are captured in the lower level of abstraction, while speech attributes such as speaker identity and emotion that persist up to one second are captured in the higher level of abstraction. This decomposition is achieved by a two-stage neural network, with a lower and an upper stage operating at different time scales. Both stages are trained to predict the content of the signal in their respective latent spaces. A top-down pathway between stages further improves the predictive capability of the network. With an application in speech compression in mind, we investigate the effect of dimensionality reduction and low bitrate quantization on the extracted representations. The performance measured on the LibriSpeech and EmoV-DB datasets reaches, and for some speech attributes even exceeds, that of state-of-the-art approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
Dimensionality ReductionQuantizationSimilar Papers 제목 키워드 기반
Relating EEG recordings to speech using envelope tracking and the speech-FFR
During speech perception, a listener's electroencephalogram (EEG) reflects acoustic-level processing as well as higher-level cognitive factors such as speech comprehension and attention. However, decoding speech from EEG…
DecoderEEGEeg DecodingElectroencephalogram (EEG)Study of cognitive component of auditory attention to natural speech events
Event-related potentials (ERP) have been used to address a wide range of research questions in neuroscience and cognitive psychology including selective auditory attention. The recent progress in auditory attention decod…
EEGElectroencephalogram (EEG)ERPAn efficient and perceptually motivated auditory neural encoding and decoding algorithm for spiking neural networks
Auditory front-end is an integral part of a spiking neural network (SNN) when performing auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reco…
Benchmarkingspeech-recognitionSpeech RecognitionSpeech segmentation with a neural encoder model of working memory
We present the first unsupervised LSTM speech segmenter as a cognitive model of the acquisition of words from unsegmented input. Cognitive biases toward phonological and syntactic predictability in speech are rooted in t…
DecoderA predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
Speech perception involves storing and integrating sequentially presented items. Recent work in cognitive neuroscience has identified temporal and contextual characteristics in humans' neural encoding of speech that may …