paper-with-me

홈 › Papers

Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language

2022-12-16 · Yusuke Yasuda, Tomoki Toda

End-to-end text-to-speech synthesis (TTS) can generate highly natural synthetic speech from raw text. However, rendering the correct pitch accents is still a challenging problem for end-to-end TTS. To tackle the challenge of rendering correct pitch accent in Japanese end-to-end TTS, we adopt PnG~BERT, a self-supervised pretrained model in the character and phoneme domain for TTS. We investigate the effects of features captured by PnG~BERT on Japanese TTS by modifying the fine-tuning condition to determine the conditions helpful inferring pitch accents. We manipulate content of PnG~BERT features from being text-oriented to speech-oriented by changing the number of fine-tuned layers during TTS. In addition, we teach PnG~BERT pitch accent information by fine-tuning with tone prediction as an additional downstream task. Our experimental results show that the features of PnG~BERT captured by pretraining contain information helpful inferring pitch accent, and PnG~BERT outperforms baseline Tacotron on accent correctness in a listening test.

📄 PDF Abstract BibTeX arXiv:2212.08321

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Sigmoid Activation 설명 없음
Highway Layer 설명 없음
BiGRU A Bidirectional GRU, or BiGRU, is a sequence processing model that consists of two GRUs. one taking the input in a forward…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Residual Connection 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Tanh Activation 설명 없음

Similar Papers 제목 키워드 기반

Benchmarking Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2

2025-05-22 · Zackary Rackauckas, Julia Hirschberg

Synthesizing expressive Japanese character speech poses unique challenges due to pitch-accent sensitivity and stylistic variability. This paper benchmarks two open-source text-to-speech models--VITS and Style-BERT-VITS2 …

BenchmarkingDialogue GenerationSensitivitytext-to-speech+1

Investigation of enhanced Tacotron text-to-speech synthesis systems with self-attention for pitch accent language

2018-10-29 · Yusuke Yasuda, Xin Wang, Shinji Takaki, Junichi Yamagishi

End-to-end speech synthesis is a promising approach that directly converts raw text to speech. Although it was shown that Tacotron2 outperforms classical pipeline systems with regards to naturalness in English, its appli…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction

2024-08-29 · Yuka Ko, Sheng Li, Chao-Han Huck Yang, Tatsuya Kawahara

With the strong representational power of large language models (LLMs), generative error correction (GER) for automatic speech recognition (ASR) aims to provide semantic and phonetic refinements to address ASR errors. Th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingLanguage Modeling+3

NAIST Simultaneous Speech Translation System for IWSLT 2024

2024-06-30 · Yuka Ko, Ryo Fukuda, Yuta Nishikawa, Yasumasa Kano 외

This paper describes NAIST's submission to the simultaneous track of the IWSLT 2024 Evaluation Campaign: English-to-{German, Japanese, Chinese} speech-to-text translation and English-to-Japanese speech-to-speech translat…

Speech-to-Speech TranslationSpeech-to-TextSpeech-to-Text Translationtext-to-speech+2

Polyphone disambiguation and accent prediction using pre-trained language models in Japanese TTS front-end

2022-01-24 · Rem Hida, Masaki Hamada, Chie Kamada, Emiru Tsunoo 외

Although end-to-end text-to-speech (TTS) models can generate natural speech, challenges still remain when it comes to estimating sentence-level phonetic and prosodic information from raw text in Japanese TTS systems. In …

Morphological AnalysisPolyphone disambiguationSentencetext-to-speech+1