paper-with-me

Papers

Benchmarking Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2

2025-05-22 · Zackary Rackauckas, Julia Hirschberg

Synthesizing expressive Japanese character speech poses unique challenges due to pitch-accent sensitivity and stylistic variability. This paper benchmarks two open-source text-to-speech models--VITS and Style-BERT-VITS2 JP Extra (SBV2JE)--on in-domain, character-driven Japanese speech. Using three character-specific datasets, we evaluate models across naturalness (mean opinion and comparative mean opinion score), intelligibility (word error rate), and speaker consistency. SBV2JE matches human ground truth in naturalness (MOS 4.37 vs. 4.38), achieves lower WER, and shows slight preference in CMOS. Enhanced by pitch-accent controls and a WavLM-based discriminator, SBV2JE proves effective for applications like language learning and character dialogue generation, despite higher computational demands.

📄 PDF Abstract BibTeX arXiv:2505.17320

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDialogue GenerationSensitivitytext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Computational Narrative Understanding for Expressive Text-to-Speech

2025-09-04 · Gaspard Michel, Elena V. Epure, Christophe Cerisara arxiv

Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data remains underexamined. We argue that human-narrated audiobooks, particularly …

Animating Language Practice: Engagement with Stylized Conversational Agents in Japanese Learning

2025-07-09 · Zackary Rackauckas, Julia Hirschberg arxiv

We explore Jouzu, a Japanese language learning application that integrates large language models with anime-inspired conversational agents. Designed to address challenges learners face in practicing natural and expressiv…

Japanese to English/Chinese/Korean Datasets for Translation Quality Estimation and Automatic Post-Editing

2017-11-01 · WS 2017 11 · Atsushi Fujita, Eiichiro Sumita

Aiming at facilitating the research on quality estimation (QE) and automatic post-editing (APE) of machine translation (MT) outputs, especially for those among Asian languages, we have created new datasets for Japanese t…

Automatic Post-EditingBenchmarkingMachine TranslationSentence+1

Benchmarking Large Language Models for Grapheme-to-Phoneme Conversion: A Japanese Case Study

2026-06-20 · Tomoki Koriyama arxiv

Grapheme-to-phoneme (G2P) conversion is essential for controllable and robust text-to-speech, and large language models (LLMs), with broad linguistic knowledge, offer a promising approach. We benchmarked over 30 LLMs on …

Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction

2024-08-29 · Yuka Ko, Sheng Li, Chao-Han Huck Yang, Tatsuya Kawahara

With the strong representational power of large language models (LLMs), generative error correction (GER) for automatic speech recognition (ASR) aims to provide semantic and phonetic refinements to address ASR errors. Th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingLanguage Modeling+3