Learning Paralinguistic Features from Audiobooks through Style Voice Conversion
Paralinguistics, the non-lexical components of speech, play a crucial role in human-human interaction. Models designed to recognize paralinguistic information, particularly speech emotion and style, are difficult to train because of the limited labeled datasets available. In this work, we present a new framework that enables a neural network to learn to extract paralinguistic attributes from speech using data that are not annotated for emotion. We assess the utility of the learned embeddings on the downstream tasks of emotion recognition and speaking style detection, demonstrating significant improvements over surface acoustic features as well as over embeddings extracted from other unsupervised approaches. Our work enables future systems to leverage the learned embedding extractor as a separate component capable of highlighting the paralinguistic components of speech.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionStyle DetectionVoice ConversionSimilar Papers 제목 키워드 기반
Evaluating expressive speech synthesis from audiobook corpora for conversational phrases
Audiobooks are a rich resource of large quantities of natural sounding, highly expressive speech. In our previous research we have shown that it is possible to detect different expressive voice styles represented in a pa…
ClusteringExpressive Speech SynthesisSpeech SynthesisInvestigating Inter- and Intra-speaker Voice Conversion using Audiobooks
Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining. A dialog is a central passage in audiobo…
Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1Large-Scale Automatic Audiobook Creation
An audiobook can dramatically improve a work of literature's accessibility and improve reader engagement. However, audiobooks can take hundreds of hours of human effort to create, edit, and publish. In this work, we pres…
text-to-speechText to SpeechStyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
This paper introduces StyleSpeech, a novel Text-to-Speech~(TTS) system that enhances the naturalness and accuracy of synthesized speech. Building upon existing TTS technologies, StyleSpeech incorporates a unique Style De…
parameter-efficient fine-tuningtext-to-speechText to SpeechMultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech an…