paper-with-me

Papers

Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis

2024-02-11 · Kenichi Fujita, Atsushi Ando, Yusuke Ijima

This paper proposes a speech rhythm-based method for speaker embeddings to model phoneme duration using a few utterances by the target speaker. Speech rhythm is one of the essential factors among speaker characteristics, along with acoustic features such as F0, for reproducing individual utterances in speech synthesis. A novel feature of the proposed method is the rhythm-based embeddings extracted from phonemes and their durations, which are known to be related to speaking rhythm. They are extracted with a speaker identification model similar to the conventional spectral feature-based one. We conducted three experiments, speaker embeddings generation, speech synthesis with generated embeddings, and embedding space analysis, to evaluate the performance. The proposed method demonstrated a moderate speaker identification performance (15.2% EER), even with only phonemes and their duration information. The objective and subjective evaluation results demonstrated that the proposed method can synthesize speech with speech rhythm closer to the target speaker than the conventional method. We also visualized the embeddings to evaluate the relationship between the distance of the embeddings and the perceptual similarity. The visualization of the embedding space and the relation analysis between the closeness indicated that the distribution of embeddings reflects the subjective and objective similarity.

📄 PDF Abstract BibTeX arXiv:2402.07085

Code (0)

등록된 구현이 없습니다.

Tasks

RhythmSpeaker IdentificationSpeech Synthesis

Similar Papers 제목 키워드 기반

Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model

2023-04-24 · Kenichi Fujita, Takanori Ashihara, Hiroki Kanagawa, Takafumi Moriya 외

This paper proposes a zero-shot text-to-speech (TTS) conditioned by a self-supervised speech-representation model acquired through self-supervised learning (SSL). Conventional methods with embedding vectors from x-vector…

RhythmSelf-Supervised LearningSpeech Synthesistext-to-speech+2

Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis

2025-07-02 · Marc-André Carbonneau, Benjamin van Niekerk, Hugo Seuté, Jean-Philippe Letendre 외 arxiv

Modeling voice identity is challenging due to its multifaceted nature. In generative speech systems, identity is often assessed using automatic speaker verification (ASV) embeddings, designed for discrimination rather th…

Speaker VerificationSpeech Synthesis

AraS2P: Arabic Speech-to-Phonemes System

2025-09-27 · Bassam Matar, Mohamed Fayed, Ayman Khalafallah arxiv

This paper describes AraS2P, our speech-to-phonemes system submitted to the Iqra'Eval 2025 Shared Task. We adapted Wav2Vec2-BERT via Two-Stage training strategy. In the first stage, task-adaptive continue pretraining was…

SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate

2022-07-13 · Nabarun Goswami, Tatsuya Harada

The mapping of text to speech (TTS) is non-deterministic, letters may be pronounced differently based on context, or phonemes can vary depending on various physiological and stylistic factors like gender, age, accent, em…

Speech Separationtext-to-speechText to Speech

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

2026-06-04 · Naman Kothari, Arjun Gangwar, Adarsh Arigala, S Umesh arxiv

Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual interference in multilingual multi-speake…