paper-with-me

홈 › Papers

Exploring attention mechanism for acoustic-based classification of speech utterances into system-directed and non-system-directed

2019-02-01 · Atta Norouzian, Bogdan Mazoure, Dermot Connolly, Daniel Willett

Voice controlled virtual assistants (VAs) are now available in smartphones, cars, and standalone devices in homes. In most cases, the user needs to first "wake-up" the VA by saying a particular word/phrase every time he or she wants the VA to do something. Eliminating the need for saying the wake-up word for every interaction could improve the user experience. This would require the VA to have the capability to detect the speech that is being directed at it and respond accordingly. In other words, the challenge is to distinguish between system-directed and non-system-directed speech utterances. In this paper, we present a number of neural network architectures for tackling this classification problem based on using only acoustic features. These architectures are based on using convolutional, recurrent and feed-forward layers. In addition, we investigate the use of an attention mechanism applied to the output of the convolutional and the recurrent layers. It is shown that incorporating the proposed attention mechanism into the models always leads to significant improvement in classification accuracy. The best model achieved equal error rates of 16.25 and 15.62 percents on two distinct realistic datasets.

📄 PDF Abstract BibTeX arXiv:1902.00570

Code (0)

등록된 구현이 없습니다.

Tasks

General Classification

Similar Papers 제목 키워드 기반

Exploring the Power of Pure Attention Mechanisms in Blind Room Parameter Estimation

2024-02-25 · Chunxi Wang, Maoshen Jia, Meiran Li, Changchun Bao 외

Dynamic parameterization of acoustic environments has drawn widespread attention in the field of audio processing. Precise representation of local room acoustic characteristics is crucial when designing audio filters for…

Data Augmentationparameter estimationTransfer Learning

Speech Emotion Recognition Using Multi-hop Attention Mechanism

2019-04-23 · 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2019 5 · Seunghyun Yoon, Seokhyun Byun, Subhadeep Dey, Kyomin Jung

In this paper, we are interested in exploiting textual and acoustic data of an utterance for the speech emotion classification task. The baseline approach models the information from audio and text independently using tw…

Emotion ClassificationEmotion RecognitionSpeech Emotion Recognition

Attention Wave-U-Net for Acoustic Echo Cancellation

2020-10-25 · Interspeech 2020 10 · Jung-Hee Kim, Joon-Hyuk Chang

In this paper, a Wave-U-Net based acoustic echo cancellation (AEC) with an attention mechanism is proposed to jointly sup- press acoustic echo and background noise. The proposed ap- proach consists of the Wave-U-Net, …

Acoustic echo cancellation

Forward Attention in Sequence-to-sequence Acoustic Modelling for Speech Synthesis

2018-07-18 · Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai

This paper proposes a forward attention method for the sequenceto- sequence acoustic modeling of speech synthesis. This method is motivated by the nature of the monotonic alignment from phone sequences to acoustic sequen…

Acoustic ModellingDecoderSpeech Synthesis

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

2026-05-25 · Loukas Ilias, Dimitris Askounis arxiv

Alzheimer's disease (AD) is a progressive neurodegenerative disorder and the leading cause of dementia, affecting memory, reasoning, communication, and daily functioning. Early diagnosis is particularly important, as tim…

Multimodal Deep LearningRepresentation Learning