paper-with-me

Papers

A Discourse-level Multi-scale Prosodic Model for Fine-grained Emotion Analysis

2023-09-21 · Xianhao Wei, Jia Jia, Xiang Li, Zhiyong Wu, Ziyi Wang

This paper explores predicting suitable prosodic features for fine-grained emotion analysis from the discourse-level text. To obtain fine-grained emotional prosodic features as predictive values for our model, we extract a phoneme-level Local Prosody Embedding sequence (LPEs) and a Global Style Embedding as prosodic speech features from the speech with the help of a style transfer model. We propose a Discourse-level Multi-scale text Prosodic Model (D-MPM) that exploits multi-scale text to predict these two prosodic features. The proposed model can be used to analyze better emotional prosodic features and thus guide the speech synthesis model to synthesize more expressive speech. To quantitatively evaluate the proposed model, we contribute a new and large-scale Discourse-level Chinese Audiobook (DCA) dataset with more than 13,000 utterances annotated sequences to evaluate the proposed model. Experimental results on the DCA dataset show that the multi-scale text information effectively helps to predict prosodic features, and the discourse-level text improves both the overall coherence and the user experience. More interestingly, although we aim at the synthesis effect of the style transfer model, the synthesized speech by the proposed text prosodic analysis model is even better than the style transfer from the original speech in some user evaluation indicators.

📄 PDF Abstract BibTeX arXiv:2309.11849

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech SynthesisStyle Transfer

Similar Papers 제목 키워드 기반

Improving Mandarin Prosodic Structure Prediction with Multi-level Contextual Information

2023-08-31 · Jie Chen, Changhe Song, Deyi Tuo, Xixin Wu 외

For text-to-speech (TTS) synthesis, prosodic structure prediction (PSP) plays an important role in producing natural and intelligible speech. Although inter-utterance linguistic information can influence the speech inter…

DecoderMulti-Task Learningtext-to-speechText to Speech

Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework

2025-10-08 · Zhu Li, Yuqing Zhang, Xiyuan Gao, Shekhar Nayak 외 arxiv

Sarcasm is a pragmatic phenomenon in which speakers convey meanings that diverge from literal content, relying on an interaction between semantics and prosodic expression. However, how these cues jointly contribute to th…

Speech Synthesis

The Prosody of Emojis

2025-08-01 · Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow arxiv

Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure. In text-based settings, where these cues are absent, emojis act as visual s…

Towards Modelling Coherence in Spoken Discourse

2020-12-31 · Rajaswa Patil, Yaman Kumar Singla, Rajiv Ratn Shah, Mika Hama 외

While there has been significant progress towards modelling coherence in written discourse, the work in modelling spoken discourse coherence has been quite limited. Unlike the coherence in text, coherence in spoken disco…

Prosodic, syntactic, semantic guidelines for topic structures across domains and corpora

2014-05-01 · LREC 2014 5 · Ana Isabel Mata, Helena Moniz, Telmo M{\'o}ia, Anabela Gon{\c{c}}alves 외

This paper presents the annotation guidelines applied to naturally occurring speech, aiming at an integrated account of contrast and parallel structures in European Portuguese. These guidelines were defined to allow for …

Part-Of-Speech Tagging