paper-with-me

Papers

SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation

2025-04-07 · Stephen Brade, Sam Anderson, Rithesh Kumar, Zeyu Jin, Anh Truong

Novice content creators often invest significant time recording expressive speech for social media videos. While recent advancements in text-to-speech (TTS) technology can generate highly realistic speech in various languages and accents, many struggle with unintuitive or overly granular TTS interfaces. We propose simplifying TTS generation by allowing users to specify high-level context alongside their script. Our Wizard-of-Oz system, SpeakEasy, leverages user-provided context to inform and influence TTS output, enabling iterative refinement with high-level feedback. This approach was informed by two 8-subject formative studies: one examining content creators' experiences with TTS, and the other drawing on effective strategies from voice actors. Our evaluation shows that participants using SpeakEasy were more successful in generating performances matching their personal standards, without requiring significantly more effort than leading industry interfaces.

📄 PDF Abstract BibTeX arXiv:2504.05106

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

SpeakEasy: A Conversational Intelligence Chatbot for Enhancing College Students' Communication Skills

2023-09-23 · Hyunbae Jeon, Rhea Ramachandran, Victoria Ploerer, Yella Diekmann 외

Social interactions and conversation skills separate the successful from the rest and the confident from the shy. For college students in particular, the ability to converse can be an outlet for the stress and anxiety ex…

Chatbot

StyleSpeech: Self-supervised Style Enhancing with VQ-VAE-based Pre-training for Expressive Audiobook Speech Synthesis

2023-12-19 · Xueyuan Chen, Xi Wang, Shaofei Zhang, Lei He 외

The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-s…

DecoderSpeech Synthesis

A Study on Altering the Latent Space of Pretrained Text to Speech Models for Improved Expressiveness

2023-11-17 · Mathias Vogel

This report explores the challenge of enhancing expressiveness control in Text-to-Speech (TTS) models by augmenting a frozen pretrained model with a Diffusion Model that is conditioned on joint semantic audio/text embedd…

text-to-speechText to Speech

Enhancing Word-Level Semantic Representation via Dependency Structure for Expressive Text-to-Speech Synthesis

2021-04-14 · Yixuan Zhou, Changhe Song, Jingbei Li, Zhiyong Wu 외

Exploiting rich linguistic information in raw text is crucial for expressive text-to-speech (TTS). As large scale pre-trained text representation develops, bidirectional encoder representations from Transformers (BERT) h…

Dependency ParsingRepresentation LearningSentenceSpeech Synthesis+3

Enhancing audio quality for expressive Neural Text-to-Speech

2021-08-13 · Abdelhamid Ezzerg, Adam Gabrys, Bartosz Putrycz, Daniel Korzekwa 외

Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recordings. However, not all speaking styles …

Acoustic ModellingSpeech Synthesistext-to-speechText to Speech