paper-with-me

Papers

Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

2026-06-22 · Jinchuan Tian, Haoran Wang, Siddhant Arora, Takashi Maekaku, Keita Goto, Jin Sakuma, Yusuke Shinohara, Chao-Han Huck Yang, Shinji Watanabe arxiv

Classical TTS systems typically rely on rigid input formats and predefined metadata slots, limiting their ability to fulfill flexible user requirements. This paper introduces Bagpiper-TTS, a universal speech synthesis system that deals with diverse natural language user requests. Given a natural language prompt, Bagpiper-TTS first reasons over the users' intent to derive a rich caption, i.e., a comprehensive textual blueprint encompassing both transcription and nuanced metadata. Subsequently, this caption guides the synthesis of the target speech. Our model inherently supports a broad spectrum of tasks besides classical TTS applications, including multi-talker, intent-to-speech, role-play synthesis, singing voice synthesis, and more. Experimental results demonstrate that Bagpiper-TTS achieves an 1.7% Word Error Rate (WER) on the Seed-TTS-Eval benchmark and match the performance of dedicated models in both LLM-as-a-judge and human subjective evaluations across multiple applications.

📄 PDF Abstract BibTeX arXiv:2606.22811

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions

2026-02-05 · Jinchuan Tian, Haoran Wang, Bo-Hao Su, Chien-yu Huang 외 arxiv

Current audio foundation models typically rely on rigid, task-specific supervision (e.g., speech recognition), addressing isolated factors of audio rather than the whole. In contrast, human processes audio holistically, …

Automatic Speech Recognition and Topic Identification for Almost-Zero-Resource Languages

2018-02-23 · Matthew Wiesner, Chunxi Liu, Lucas Ondel, Craig Harman 외

Automatic speech recognition (ASR) systems often need to be developed for extremely low-resource languages to serve end-uses such as audio content categorization and search. While universal phone recognition is natural t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Humanitarianspeech-recognition+1

Tools and resources for Romanian text-to-speech and speech-to-text applications

2018-02-15 · Tiberiu Boros, Stefan Daniel Dumitrescu, Vasile Pais

In this paper we introduce a set of resources and tools aimed at providing support for natural language processing, text-to-speech synthesis and speech recognition for Romanian. While the tools are general purpose and ca…

speech-recognitionSpeech RecognitionSpeech SynthesisSpeech-to-Text+3

On The Landscape of Spoken Language Models: A Comprehensive Survey

2025-04-11 · Siddhant Arora, Kai-Wei Chang, Chung-Ming Chien, Yifan Peng 외

The field of spoken language processing is undergoing a shift from training custom-built, task-specific models toward using and optimizing spoken language models (SLMs) which act as universal speech processing systems. T…

Survey

Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS

2023-09-14 · Yifan Yang, Feiyu Shen, Chenpeng Du, Ziyang Ma 외

Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer lower storage requirements and great po…

Self-Supervised Learningspeech-recognitionSpeech RecognitionSpeech Synthesis