paper-with-me

홈 › Papers

Continuous Speech for Improved Learning Pathological Voice Disorders

2022-02-22 · Syu-Siang Wang, Chi-Te Wang, Chih-Chung Lai, Yu Tsao, Shih-Hau Fang

Goal: Numerous studies had successfully differentiated normal and abnormal voice samples. Nevertheless, further classification had rarely been attempted. This study proposes a novel approach, using continuous Mandarin speech instead of a single vowel, to classify four common voice disorders (i.e. functional dysphonia, neoplasm, phonotrauma, and vocal palsy). Methods: In the proposed framework, acoustic signals are transformed into mel-frequency cepstral coefficients, and a bi-directional long-short term memory network (BiLSTM) is adopted to model the sequential features. The experiments were conducted on a large-scale database, wherein 1,045 continuous speech were collected by the speech clinic of a hospital from 2012 to 2019. Results: Experimental results demonstrated that the proposed framework yields significant accuracy and unweighted average recall improvements of 78.12-89.27% and 50.92-80.68%, respectively, compared with systems that use a single vowel. Conclusions: The results are consistent with other machine learning algorithms, including gated recurrent units, random forest, deep neural networks, and LSTM. The sensitivities for each disorder were also analyzed, and the model capabilities were visualized via principal component analysis. An alternative experiment based on a balanced dataset again confirms the advantages of using continuous speech for learning voice disorders.

📄 PDF Abstract BibTeX arXiv:2202.10777

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…
Memory Network 설명 없음

Similar Papers 제목 키워드 기반

Voice Disorder Detection Using Long Short Term Memory (LSTM) Model

2018-12-04 · Vibhuti Gupta

Automated detection of voice disorders with computational methods is a recent research area in the medical domain since it requires a rigorous endoscopy for the accurate diagnosis. Efficient screening methods are require…

Specificity

Robustness against the channel effect in pathological voice detection

2018-11-26 · Yi-Te Hsu, Zining Zhu, Chi-Te Wang, Shih-Hau Fang 외

Many people are suffering from voice disorders, which can adversely affect the quality of their lives. In response, some researchers have proposed algorithms for automatic assessment of these disorders, based on voice si…

Domain AdaptationUnsupervised Domain Adaptation

Deep Learning for Pathological Speech: A Survey

2025-01-07 · Shakeel A. Sheikh, Md. Sahidullah, Ina Kodrasi

Advancements in spoken language technologies for neurodegenerative speech disorders are crucial for meeting both clinical and technological needs. This overview paper is vital for advancing the field, as it presents a co…

Automatic Speech RecognitionData AugmentationDeep Learningspeech-recognition+2

AI-Driven Acoustic Voice Biomarker-Based Hierarchical Classification of Benign Laryngeal Voice Disorders from Sustained Vowels

2025-12-31 · Mohsen Annabestani, Samira Aghadoost, Anais Rameau, Olivier Elemento 외 arxiv

Benign laryngeal voice disorders affect nearly one in five individuals and often manifest as dysphonia, while also serving as non-invasive indicators of broader physiological dysfunction. We introduce a clinically inspir…

Unified Pathological Speech Analysis with Prompt Tuning

2024-11-05 · Fei Yang, Xuenan Xu, Mengyue Wu, Kai Yu

Pathological speech analysis has been of interest in the detection of certain diseases like depression and Alzheimer's disease and attracts much interest from researchers. However, previous pathological speech analysis m…

Language ModelingLanguage Modelling