paper-with-me

홈 › Papers

A Semi-Supervised Framework for Speech Confidence Detection using Whisper

2026-05-12 · Adam Wynn, Jingyun Wang arxiv

Automatic detection of speaker confidence is critical for adaptive computing but remains constrained by limited labelled data and the subjectivity of paralinguistic annotations. This paper proposes a semi-supervised hybrid framework that fuses deep semantic embeddings from the Whisper encoder with an interpretable acoustic feature vector composed of eGeMAPS descriptors and auxiliary probability estimates of vocal stress and disfluency. To mitigate reliance on scarce ground truth data, we introduce an Uncertainty-Aware Pseudo-Labelling strategy where a model generates labels for unlabelled data, retaining only high-quality samples for training. Experimental results demonstrate that the proposed approach achieves a Macro-F1 score of 0.751, outperforming self-supervised baselines, including WavLM, HuBERT, and Wav2Vec 2.0. The hybrid architecture also surpasses the unimodal Whisper baseline, yielding a 3\% improvement in the minority class, confirming that explicit prosodic and auxiliary features provide necessary corrective signals which are otherwise lost in deep semantic representations. Ablation studies further show that a curated set of high confidence pseudo-labels outperforms indiscriminate large scale augmentation, confirming that data quality outweighs quantity for perceived confidence detection.

📄 PDF Abstract BibTeX arXiv:2605.12387

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling

2024-09-25 · Yuanchao Li, Zixing Zhang, Jing Han, Peter Bell 외

The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervi…

Automatic Speech RecognitionEmotion Recognitionspeech-recognitionSpeech Recognition

Semi-Supervised Speech Confidence Detection using Pseudo-Labelling and Whisper Embeddings

2026-06-15 · Adam Wynn, Jingyun Wang, Xiangyu Tan arxiv

Understanding speaker confidence is crucial in educational settings, as it can enhance personalised feedback and improve learning outcomes. This study introduces a novel framework for detecting speaker confidence by inte…

Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition

2023-08-12 · Han Zhu, Dongji Gao, Gaofeng Cheng, Daniel Povey 외

When labeled data is insufficient, semi-supervised learning with the pseudo-labeling technique can significantly improve the performance of automatic speech recognition. However, pseudo-labels are often noisy, containing…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Weakly-supervised text-to-speech alignment confidence measure

2016-12-01 · COLING 2016 12 · Guillaume Serri{\`e}re, Christophe Cerisara, Dominique Fohr, Odile Mella

This work proposes a new confidence measure for evaluating text-to-speech alignment systems outputs, which is a key component for many applications, such as semi-automatic corpus anonymization, lips syncing, film dubbing…

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+1

Using heterogeneity in semi-supervised transcription hypotheses to improve code-switched speech recognition

2021-06-14 · Andrew Slottje, Shannon Wotherspoon, William Hartmann, Matthew Snover 외

Modeling code-switched speech is an important problem in automatic speech recognition (ASR). Labeled code-switched data are rare, so monolingual data are often used to model code-switched speech. These monolingual data m…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition