paper-with-me

Papers

Selfsupervised learning for pathological speech detection

2024-05-16 · Shakeel Ahmad Sheikh

Speech production is a complex phenomenon, wherein the brain orchestrates a sequence of processes involving thought processing, motor planning, and the execution of articulatory movements. However, this intricate execution of various processes is susceptible to influence and disruption by various neurodegenerative pathological speech disorders, such as Parkinsons' disease, resulting in dysarthria, apraxia, and other conditions. These disorders lead to pathological speech characterized by abnormal speech patterns and imprecise articulation. Diagnosing these speech disorders in clinical settings typically involves auditory perceptual tests, which are time-consuming, and the diagnosis can vary among clinicians based on their experiences, biases, and cognitive load during the diagnosis. Additionally, unlike neurotypical speakers, patients with speech pathologies or impairments are unable to access various virtual assistants such as Alexa, Siri, etc. To address these challenges, several automatic pathological speech detection (PSD) approaches have been proposed. These approaches aim to provide efficient and accurate detection of speech disorders, thereby facilitating timely intervention and support for individuals affected by these conditions. These approaches mainly vary in two aspects: the input representations utilized and the classifiers employed. Due to the limited availability of data, the performance of detection remains subpar. Self-supervised learning (SSL) embeddings, such as wav2vec2, and their multilingual versions, are being explored as a promising avenue to improve performance. These embeddings leverage self-supervised learning techniques to extract rich representations from audio data, thereby offering a potential solution to address the limitations posed by the scarcity of labeled data.

📄 PDF Abstract BibTeX arXiv:2406.02572

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

Impact of Speech Mode in Automatic Pathological Speech Detection

2024-06-14 · Shakeel A. Sheikh, Ina Kodrasi

Automatic pathological speech detection approaches yield promising results in identifying various pathologies. These approaches are typically designed and evaluated for phonetically-controlled speech scenarios, where spe…

Navigate

Glottal Closure Instants Detection From Pathological Acoustic Speech Signal Using Deep Learning

2018-11-25 · Gurunath Reddy M, Tanumay Mandal, Krothapalli Sreenivasa Rao

In this paper, we propose a classification based glottal closure instants (GCI) detection from pathological acoustic speech signal, which finds many applications in vocal disorder analysis. Till date, GCI for pathologica…

General Classification

Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform

2025-08-14 · Yuankun Xie, Ruibo Fu, Xiaopeng Wang, Zhiyong Wang 외 arxiv

The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on publ…

DeepFake DetectionData Augmentation

Multiview Canonical Correlation Analysis for Automatic Pathological Speech Detection

2024-09-13 · Yacouba Kaloga, Shakeel A. Sheikh, Ina Kodrasi

Recently proposed automatic pathological speech detection approaches rely on spectrogram input representations or wav2vec2 embeddings. These representations may contain pathology irrelevant uncorrelated information, such…

Dimensionality Reduction

Exploring In-Context Learning Capabilities of ChatGPT for Pathological Speech Detection

2025-03-31 · Mahdi Amiri, Hatef Otroshi Shahreza, Ina Kodrasi

Automatic pathological speech detection approaches have shown promising results, gaining attention as potential diagnostic tools alongside costly traditional methods. While these approaches can achieve high accuracy, the…

DiagnosticIn-Context Learning