paper-with-me

Papers

Bridging the prosody GAP: Genetic Algorithm with People to efficiently sample emotional prosody

2022-05-10 · Pol van Rijn, Harin Lee, Nori Jacoby

The human voice effectively communicates a range of emotions with nuanced variations in acoustics. Existing emotional speech corpora are limited in that they are either (a) highly curated to induce specific emotions with predefined categories that may not capture the full extent of emotional experiences, or (b) entangled in their semantic and prosodic cues, limiting the ability to study these cues separately. To overcome this challenge, we propose a new approach called 'Genetic Algorithm with People' (GAP), which integrates human decision and production into a genetic algorithm. In our design, we allow creators and raters to jointly optimize the emotional prosody over generations. We demonstrate that GAP can efficiently sample from the emotional speech space and capture a broad range of emotions, and show comparable results to state-of-the-art emotional speech corpora. GAP is language-independent and supports large crowd-sourcing, thus can support future large-scale cross-cultural research.

📄 PDF Abstract BibTeX arXiv:2205.04820

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Prosody: Models, Methods, and Applications

2021-08-01 · ACL 2021 5 · Nigel Ward, Gina-Anne Levow

Prosody is essential in human interaction, enabling people to show interest, establish rapport, efficiently convey nuances of attitude or intent, and so on. Some applications that exploit prosodic knowledge have recently…

Pitchtron: Towards audiobook generation from ordinary people's voices

2020-05-21 · Interspeech 2020 5 · Sunghee Jung, Hoirin Kim

In this paper, we explore prosody transfer for audiobook generation under rather realistic condition where training DB is plain audio mostly from multiple ordinary people and reference audio given during inference is fro…

Decoder

Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models

2025-06-01 · Kyowoon Lee, Artyom Stitsyuk, Gunu Jho, Inchul Hwang 외

Recent advances in Text-to-Speech (TTS) have significantly improved speech naturalness, increasing the demand for precise prosody control and mispronunciation correction. Existing approaches for prosody manipulation ofte…

counterfactualSpeech Synthesistext-to-speechText to Speech

Computation of Diet Composition for Patients Suffering from Kidney and Urinary Tract Diseases with the Fuzzy Genetic System

2013-06-25 · Sri Hartati, Shofwatul 'Uyun

Determination of dietary food consumed a day for patients with diseases in general, greatly affect the health of the body and the healing process, is no exception for people with kidney disease and urinary tract. This pa…

Prosody Analysis of Audiobooks

2023-10-10 · Charuta Pethe, Bach Pham, Felix D Childress, Yunting Yin 외

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…

AttributeLanguage ModelingLanguage ModellingProsody Prediction+2