paper-with-me

홈 › Papers

Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling

2024-09-25 · Yuanchao Li, Zixing Zhang, Jing Han, Peter Bell, Catherine Lai

The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling method that leverages both acoustic and linguistic characteristics to select the most confident data for training the classification model. Acoustically, unlabeled data are compared to labeled data using the Frechet audio distance, calculated from embeddings generated by multiple audio encoders. Linguistically, large language models are prompted to revise automatic speech recognition transcriptions and predict labels based on our proposed task-specific knowledge. High-confidence data are identified when pseudo-labels from both sources align, while mismatches are treated as low-confidence data. A bimodal classifier is then trained to iteratively label the low-confidence data until a predefined criterion is met. We evaluate our SSL framework on emotion recognition and dementia detection tasks. Experimental results demonstrate that our method achieves competitive performance compared to fully supervised learning using only 30% of the labeled data and significantly outperforms two selected baselines.

📄 PDF Abstract BibTeX arXiv:2409.16937

Code (1)

yc-li20/semi-supervised-training 공식 구현 pytorch

Tasks

Automatic Speech RecognitionEmotion Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Hypergraph based semi-supervised learning algorithms applied to speech recognition problem: a novel approach

2018-10-28 · Loc Hoang Tran, Trang Hoang, Bui Hoang Nam Huynh

Most network-based speech recognition methods are based on the assumption that the labels of two adjacent speech samples in the network are likely to be the same. However, assuming the pairwise relationship between speec…

Sensitivityspeech-recognitionSpeech Recognition

Label Propagation-Based Semi-Supervised Learning for Hate Speech Classification

2020-11-01 · EMNLP (insights) 2020 11 · Ashwin Geet D’Sa, Irina Illina, Dominique Fohr, Dietrich Klakow 외

Research on hate speech classification has received increased attention. In real-life scenarios, a small amount of labeled hate speech data is available to train a reliable classifier. Semi-supervised learning takes adva…

Classification

Semi-supervised Learning with Sparse Autoencoders in Phone Classification

2016-10-03 · Akash Kumar Dhaka, Giampiero Salvi

We propose the application of a semi-supervised learning method to improve the performance of acoustic modelling for automatic speech recognition based on deep neural net- works. As opposed to unsupervised initialisation…

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Classification+3

Semi-supervised Acoustic Modelling for Five-lingual Code-switched ASR using Automatically-segmented Soap Opera Speech

2020-05-01 · LREC 2020 5 · Nick Wilkinson, Astik Biswas, Emre Yilmaz, Febe De Wet 외

This paper considers the impact of automatic segmentation on the fully-automatic, semi-supervised training of automatic speech recog-nition (ASR) systems for five-lingual code-switched (CS) speech. Four automatic segment…

Acoustic ModellingAction DetectionActivity DetectionSegmentation+2

Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

2026-05-26 · Serli Kopar, Roshan Prakash Rane, Christian Mychajliw, Lydia Federmann 외 arxiv

This study examines the relationship between speech representations and the hierarchical structure of cognitive assessment in mild cognitive impairment. Utilizing 5,754 German neuropsychological assessment recordings, we…

Self-Supervised Learning