paper-with-me

홈 › Papers

Residual Attention Based Network for Automatic Classification of Phonation Modes

2021-07-18 · Xiaoheng Sun, Yiliang Jiang, Wei Li

Phonation mode is an essential characteristic of singing style as well as an important expression of performance. It can be classified into four categories, called neutral, breathy, pressed and flow. Previous studies used voice quality features and feature engineering for classification. While deep learning has achieved significant progress in other fields of music information retrieval (MIR), there are few attempts in the classification of phonation modes. In this study, a Residual Attention based network is proposed for automatic classification of phonation modes. The network consists of a convolutional network performing feature processing and a soft mask branch enabling the network focus on a specific area. In comparison experiments, the models with proposed network achieve better results in three of the four datasets than previous works, among which the highest classification accuracy is 94.58%, 2.29% higher than the baseline.

📄 PDF Abstract BibTeX arXiv:2107.08425

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationFeature EngineeringInformation RetrievalMusic Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models

2026-02-14 · Aju Ani Justus, Ruchit Agrawal, Sudarsana Reddy Kadiri, Shrikanth Narayanan arxiv

We present voice2mode, a method for classification of four singing phonation modes (breathy, neutral (modal), flow, and pressed) using embeddings extracted from large self-supervised speech models. Prior work on singing …

Speech Recognition

Classification of ALS patients based on acoustic analysis of sustained vowel phonations

2020-12-14 · Maxim Vashkevich, Yulia Rushkevich

Amyotrophic lateral sclerosis (ALS) is incurable neurological disorder with rapidly progressive course. Common early symptoms of ALS are difficulty in swallowing and speech. However, early acoustic manifestation of speec…

feature selectionGeneral ClassificationSensitivitySpecificity

Towards detecting the pathological subharmonic voicing with fully convolutional neural networks

2025-01-15 · Takeshi Ikuma, Melda Kunduk, Brad Story, Andrew J. McWhorter

Many voice disorders induce subharmonic phonation, but voice signal analysis is currently lacking a technique to detect the presence of subharmonics reliably. Distinguishing subharmonic phonation from normal phonation is…

Interpreting glottal flow dynamics for detecting COVID-19 from voice

2020-10-29 · Soham Deshmukh, Mahmoud Al Ismail, Rita Singh

In the pathogenesis of COVID-19, impairment of respiratory functions is often one of the key symptoms. Studies show that in these cases, voice production is also adversely affected -- vocal fold oscillations are asynchro…

A Database of Laryngeal High-Speed Videos with Simultaneous High-Quality Audio Recordings of Pathological and Non-Pathological Voices

2016-05-01 · LREC 2016 5 · Philipp Aichinger, Immer Roesner, Matthias Leonhard, Doris-Maria Denk-Linnert 외

Auditory voice quality judgements are used intensively for the clinical assessment of pathological voice. Voice quality concepts are fuzzily defined and poorly standardized however, which hinders scientific and clinical …

Vocal Bursts Intensity Prediction