paper-with-me

홈 › Papers

MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data

2026-06-16 · Subhankar Ghosh, Jason Li, Paarth Neekhara, Shehzeen Hussain, Ryan Langman, Xuesong Yang, Roy Fejgin arxiv

Neural Text-to-Speech (TTS) systems achieve remarkable quality on short utterances but long-form speech generation shows prosodic drift, speaker inconsistencies and sentence boundary artifacts. Existing approaches either compress sequences, increase context length or naively concatenate independently synthesized chunks. We present an inference-time approach called MagpieTTS-LF that enables MagpieTTS to produce coherent long-form speech without model retraining. Our method introduces three key innovations: (1) soft attention priors to guide monotonic alignment while preserving past and future context; (2) a stateful inference algorithm that maintains context across sentence chunks, ensuring prosodic continuity; (3) history-aware text encoding that uses past text for discourse-level prosodic planning. Experiments on long texts show significant improvements in long-range intelligibility, prosodic coherence, speaker consistency, and boundary naturalness compared to other baselines.

📄 PDF Abstract BibTeX arXiv:2606.18485

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Long-Form Speech Generation with Spoken Language Models

2024-12-24 · Se Jin Park, Julian Salazar, Aren Jansen, Keisuke Kinoshita 외

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, current spoken language models struggle to generate plaus…

FormLanguage ModelingLanguage Modelling

Accelerating Transducers through Adjacent Token Merging

2023-06-28 · Yuang Li, Yu Wu, Jinyu Li, Shujie Liu

Recent end-to-end automatic speech recognition (ASR) systems often utilize a Transformer-based acoustic encoder that generates embedding at a high frame rate. However, this design is inefficient, particularly for long sp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)GPUspeech-recognition+1

VADOI:Voice-Activity-Detection Overlapping Inference For End-to-end Long-form Speech Recognition

2022-02-22 · Jinhan Wang, Xiaosu Tong, Jinxi Guo, Di He 외

While end-to-end models have shown great success on the Automatic Speech Recognition task, performance degrades severely when target sentences are long-form. The previous proposed methods, (partial) overlapping inference…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+3

Online Binaural Speech Separation of Moving Speakers With a Wavesplit Network

2023-03-13 · Cong Han, Nima Mesgarani

Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT) for training. In inference time, howeve…

Online ClusteringSpeaker SeparationSpeech Separation

Deep Feed-forward Sequential Memory Networks for Speech Synthesis

2018-02-26 · Mengxiao Bi, Heng Lu, Shiliang Zhang, Ming Lei 외

The Bidirectional LSTM (BLSTM) RNN based speech synthesis system is among the best parametric Text-to-Speech (TTS) systems in terms of the naturalness of generated speech, especially the naturalness in prosody. However, …

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+1