paper-with-me

홈 › Papers

Read it to me: An emotionally aware Speech Narration Application

2022-09-06 · Rishibha Bansal

In this work we try to perform emotional style transfer on audios. In particular, MelGAN-VC architecture is explored for various emotion-pair transfers. The generated audio is then classified using an LSTM-based emotion classifier for audio. We find that "sad" audio is generated well as compared to "happy" or "anger" as people have similar expressions of sadness.

📄 PDF Abstract BibTeX arXiv:2209.02785

Code (0)

등록된 구현이 없습니다.

Tasks

Style Transfer

Similar Papers 제목 키워드 기반

Making Social Platforms Accessible: Emotion-Aware Speech Generation with Integrated Text Analysis

2024-10-24 · Suparna De, Ionut Bostan, Nishanth Sastry

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to…

Speech Synthesistext-to-speechText to Speech

Reading the Mood Behind Words: Integrating Prosody-Derived Emotional Context into Socially Responsive VR Agents

2026-03-10 · SangYeop Jeong, Yeongseo Na, Seung Gyu Jeong, Jin-Woo Jeong 외 arxiv

In VR interactions with embodied conversational agents, users' emotional intent is often conveyed more by how something is said than by what is said. However, most VR agent pipelines rely on speech-to-text processing, di…

Speech Emotion Recognition

Prosody Analysis of Audiobooks

2023-10-10 · Charuta Pethe, Bach Pham, Felix D Childress, Yunting Yin 외

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…

AttributeLanguage ModelingLanguage ModellingProsody Prediction+2

Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles

2025-10-03 · Rongchen Guo, Vincent Francoeur, Isar Nejadgholi, Sylvain Gagnon 외 arxiv

Speech Emotion Recognition (SER) is essential for improving human-computer interaction, yet its accuracy remains constrained by the complexity of emotional nuances in speech. In this study, we distinguish between descrip…

Speech Emotion Recognition

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning

2026-05-30 · Sukru Samet Dindar, Riki Shimizu, Xilin Jiang, Nima Mesgarani arxiv

Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues. To address this, we introduce Sympatheia, a s…