A Novel Scheme to classify Read and Spontaneous Speech
The COVID-19 pandemic has led to an increased use of remote telephonic interviews, making it important to distinguish between scripted and spontaneous speech in audio recordings. In this paper, we propose a novel scheme for identifying read and spontaneous speech. Our approach uses a pre-trained DeepSpeech audio-to-alphabet recognition engine to generate a sequence of alphabets from the audio. From these alphabets, we derive features that allow us to discriminate between read and spontaneous speech. Our experimental results show that even a small set of self-explanatory features can effectively classify the two types of speech very effectively.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
On the Use of Self-Supervised Speech Representations in Spontaneous Speech Synthesis
Self-supervised learning (SSL) speech representations learned from large amounts of diverse, mixed-quality speech data without transcriptions are gaining ground in many speech technology applications. Prior work has show…
PredictionSelf-Supervised LearningSpeech Synthesistext-to-speech+1Automatic Anomaly Detection for Dysarthria across Two Speech Styles: Read vs Spontaneous Speech
Perceptive evaluation of speech disorders is still the standard method in clinical practice for the diagnosing and the following of the condition progression of patients. Such methods include different tasks such as read…
Anomaly DetectionAdaSpeech 3: Adaptive Text to Speech for Spontaneous Style
While recent text to speech (TTS) models perform very well in synthesizing reading-style (e.g., audiobook) speech, it is still challenging to synthesize spontaneous-style speech (e.g., podcast or conversation), mainly be…
DecoderMixture-of-ExpertsRhythmtext-to-speech+1SpeechYOLO: Detection and Localization of Speech Objects
In this paper, we propose to apply object detection methods from the vision domain on the speech recognition domain, by treating audio fragments as objects. More specifically, we present SpeechYOLO, which is inspired by …
General ClassificationKeyword SpottingObjectobject-detection+3Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis
The spontaneous behavior that often occurs in conversations makes speech more human-like compared to reading-style. However, synthesizing spontaneous-style speech is challenging due to the lack of high-quality spontaneou…
Expressive Speech SynthesisSentenceSpeech Synthesistext-to-speech+2