paper-with-me

Papers

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers

2025-05-19 · Kyeongman Park, Seongho Joo, Kyomin Jung

We introduce MultiActor-Audiobook, a zero-shot approach for generating audiobooks that automatically produces consistent, expressive, and speaker-appropriate prosody, including intonation and emotion. Previous audiobook systems have several limitations: they require users to manually configure the speaker's prosody, read each sentence with a monotonic tone compared to voice actors, or rely on costly training. However, our MultiActor-Audiobook addresses these issues by introducing two novel processes: (1) MSP (Multimodal Speaker Persona Generation) and (2) LSI (LLM-based Script Instruction Generation). With these two processes, MultiActor-Audiobook can generate more emotionally expressive audiobooks with a consistent speaker prosody without additional training. We compare our system with commercial products, through human and MLLM evaluations, achieving competitive results. Furthermore, we demonstrate the effectiveness of MSP and LSI through ablation studies.

📄 PDF Abstract BibTeX arXiv:2505.13082

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Zero-Shot Text-to-Speech for Vietnamese

2025-06-02 · Thi Vu, Linh The Nguyen, Dat Quoc Nguyen

This paper introduces PhoAudiobook, a newly curated dataset comprising 941 hours of high-quality audio for Vietnamese text-to-speech. Using PhoAudiobook, we conduct experiments on three leading zero-shot TTS models: VALL…

text-to-speechText to Speech

Large-Scale Automatic Audiobook Creation

2023-09-07 · Brendan Walsh, Mark Hamilton, Greg Newby, Xi Wang 외

An audiobook can dramatically improve a work of literature's accessibility and improve reader engagement. However, audiobooks can take hundreds of hours of human effort to create, edit, and publish. In this work, we pres…

text-to-speechText to Speech

ReadAlong Studio: Practical Zero-Shot Text-Speech Alignment for Indigenous Language Audiobooks

2022-06-01 · SIGUL (LREC) 2022 6 · Patrick Littell, Eric Joanis, Aidan Pine, Marc Tessier 외

While the alignment of audio recordings and text (often termed “forced alignment”) is often treated as a solved problem, in practice the process of adapting an alignment system to a new, under-resourced language comes wi…

HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset

2025-06-04 · Ryan Langman, Xuesong Yang, Paarth Neekhara, Shehzeen Hussain 외

This paper introduces HiFiTTS-2, a large-scale speech dataset designed for high-bandwidth speech synthesis. The dataset is derived from LibriVox audiobooks, and contains approximately 36.7k hours of English speech for 22…

Speech Synthesistext-to-speechText to Speech

Prosody Analysis of Audiobooks

2023-10-10 · Charuta Pethe, Bach Pham, Felix D Childress, Yunting Yin 외

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…

AttributeLanguage ModelingLanguage ModellingProsody Prediction+2