paper-with-me

홈 › Papers

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model

2025-09-01 · Joonyong Park, Daisuke Saito, Nobuaki Minematsu arxiv

This study presents a novel approach to voice synthesis that can substitute the traditional grapheme-to-phoneme (G2P) conversion by using a deep learning-based model that generates discrete tokens directly from speech. Utilizing a pre-trained voice SSL model, we train a T5 encoder to produce pseudo-language labels from mixed-script texts (e.g., containing Kanji and Kana). This method eliminates the need for manual phonetic transcription, reducing costs and enhancing scalability, especially for large non-transcribed audio datasets. Our model matches the performance of conventional G2P-based text-to-speech systems and is capable of synthesizing speech that retains natural linguistic and paralinguistic features, such as accents and intonations.

📄 PDF Abstract BibTeX arXiv:2509.01391

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSpeech Synthesis

Similar Papers 제목 키워드 기반

MixedGaussianAvatar: Realistically and Geometrically Accurate Head Avatar via Mixed 2D-3D Gaussian Splatting

2024-12-06 · Peng Chen, Xiaobao Wei, Qingpo Wuwu, Xinyi Wang 외

Reconstructing high-fidelity 3D head avatars is crucial in various applications such as virtual reality. The pioneering methods reconstruct realistic head avatars with Neural Radiance Fields (NeRF), which have been limit…

3DGSNeRF

JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

2017-10-28 · Ryosuke Sonobe, Shinnosuke Takamichi, Hiroshi Saruwatari

Thanks to improvements in machine learning techniques including deep learning, a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role. However, su…

BIG-bench Machine LearningSpeech Synthesis

MixedGrad: An O(1/T) Convergence Rate Algorithm for Stochastic Smooth Optimization

2013-07-26 · Mehrdad Mahdavi, Rong Jin

It is well known that the optimal convergence rate for stochastic optimization of smooth functions is $O(1/\sqrt{T})$, which is same as stochastic optimization of Lipschitz continuous convex functions. This is in contras…

Stochastic Optimization

Speech Synthesis of Code-Mixed Text

2016-05-01 · LREC 2016 5 · Sunayana Sitaram, Alan W. black

Most Text to Speech (TTS) systems today assume that the input text is in a single language and is written in the same language that the text needs to be synthesized in. However, in bilingual and multilingual communities,…

Language IdentificationSpeech Synthesistext-to-speechText to Speech

SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution

2025-10-27 · Dharma Teja Donepudi arxiv

Intra-sentence multilingual speech synthesis (code-switching TTS) remains a major challenge due to abrupt language shifts, varied scripts, and mismatched prosody between languages. Conventional TTS systems are typically …

Language IdentificationSpeech Synthesis