paper-with-me

홈 › Papers

Augmenting Bottleneck Features of Deep Neural Network Employing Motor State for Speech Recognition at Humanoid Robots

2018-08-27 · Moa Lee, Joon Hyuk Chang

As for the humanoid robots, the internal noise, which is generated by motors, fans and mechanical components when the robot is moving or shaking its body, severely degrades the performance of the speech recognition accuracy. In this paper, a novel speech recognition system robust to ego-noise for humanoid robots is proposed, in which on/off state of the motor is employed as auxiliary information for finding the relevant input features. For this, we consider the bottleneck features, which have been successfully applied to deep neural network (DNN) based automatic speech recognition (ASR) system. When learning the bottleneck features to catch, we first exploit the motor on/off state data as supplementary information in addition to the acoustic features as the input of the first deep neural network (DNN) for preliminary acoustic modeling. Then, the second DNN for primary acoustic modeling employs both the bottleneck features tossed from the first DNN and the acoustics features. When the proposed method is evaluated in terms of phoneme error rate (PER) on TIMIT database, the experimental results show that achieve obvious improvement (11% relative) is achieved by our algorithm over the conventional systems.

📄 PDF Abstract BibTeX arXiv:1808.08702

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Online set-point estimation for feedback-based traffic control applications

2022-07-27 · Farzam Tajdari, Claudio Roncoli

This paper deals with traffic control at motorway bottlenecks assuming the existence of an unknown, time-varying, Fundamental Diagram (FD). The FD may change over time due to different traffic compositions, e.g., light a…

AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning

2026-02-14 · Siqian Tong, Xuan Li, Yiwei Wang, Baolong Bi 외 arxiv

Large Audio Language Models (LALMs) excel at perception but struggle with complex reasoning requiring precise acoustic measurements. While external tools can extract fine-grained features like exact tempo or pitch, effec…

Reinforcement Learning

Exploration of Various Fractional Order Derivatives in Parkinson's Disease Dysgraphia Analysis

2023-01-20 · Jan Mucha, Zoltan Galaz, Jiri Mekyska, Marcos Faundez-Zanuy 외

Parkinson's disease (PD) is a common neurodegenerative disorder with a prevalence rate estimated to 2.0% for people aged over 65 years. Cardinal motor symptoms of PD such as rigidity and bradykinesia affect the muscles i…

Specificity

PreMovNet: Pre-Movement EEG-based Hand Kinematics Estimation for Grasp and Lift task

2022-05-02 · Anant Jain, Lalan Kumar

Kinematics decoding from brain activity helps in developing rehabilitation or power-augmenting brain-computer interface devices. Low-frequency signals recorded from non-invasive electroencephalography (EEG) are associate…

Brain Computer InterfaceEEGElectroencephalogram (EEG)

Pseudo-Inverted Bottleneck Convolution for DARTS Search Space

2022-12-31 · Arash Ahmadian, Louis S. P. Liu, Yue Fei, Konstantinos N. Plataniotis 외

Differentiable Architecture Search (DARTS) has attracted considerable attention as a gradient-based neural architecture search method. Since the introduction of DARTS, there has been little work done on adapting the acti…

Neural Architecture Search