paper-with-me

Papers

Personalized Voice Synthesis through Human-in-the-Loop Coordinate Descent

2024-08-30 · Yusheng Tian, Junbin Liu, Tan Lee

This paper describes a human-in-the-loop approach to personalized voice synthesis in the absence of reference speech data from the target speaker. It is intended to help vocally disabled individuals restore their lost voices without requiring any prior recordings. The proposed approach leverages a learned speaker embedding space. Starting from an initial voice, users iteratively refine the speaker embedding parameters through a coordinate descent-like process, guided by auditory perception. By analyzing the latent space, it is noted that that the embedding parameters correspond to perceptual voice attributes, including pitch, vocal tension, brightness, and nasality, making the search process intuitive. Computer simulations and real-world user studies demonstrate that the proposed approach is effective in approximating target voices across a diverse range of test cases.

📄 PDF Abstract BibTeX arXiv:2408.17068

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Empirical Study Incorporating Linguistic Knowledge on Filled Pauses for Personalized Spontaneous Speech Synthesis

2022-10-14 · Yuta Matsunaga, Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari

We present a comprehensive empirical study for personalized spontaneous speech synthesis on the basis of linguistic knowledge. With the advent of voice cloning for reading-style speech synthesis, a new voice cloning para…

Speech SynthesisVoice Cloning

FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning

2026-01-16 · Tanyu Chen, Tairan Chen, Kai Shen, Zhenghua Bao 외 arxiv

Recent end-to-end spoken dialogue systems leverage speech tokenizers and neural audio codecs to enable LLMs to operate directly on discrete speech representations. However, these models often exhibit limited speaker iden…

PresentCoach: Dual-Agent Presentation Coaching through Exemplars and Interactive Feedback

2025-11-19 · Sirui Chen, Jinsong Zhou, Xinli Xu, Xiaoyu Yang 외 arxiv

Effective presentation skills are essential in education, professional communication, and public speaking, yet learners often lack access to high-quality exemplars or personalized coaching. Existing AI tools typically pr…

Zero-shot personalized lip-to-speech synthesis with face image based voice control

2023-05-09 · Zheng-Yan Sheng, Yang Ai, Zhen-Hua Ling

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. Howev…

Lip to Speech SynthesisRepresentation LearningSpeech Synthesis

Creating Personalized Synthetic Voices from Post-Glossectomy Speech with Guided Diffusion Models

2023-05-27 · Yusheng Tian, Guangyan Zhang, Tan Lee

This paper is about developing personalized speech synthesis systems with recordings of mildly impaired speech. In particular, we consider consonant and vowel alterations resulted from partial glossectomy, the surgical r…

Speech SynthesisVoice Conversion