paper-with-me

Papers

P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech Prompting

2023-09-22 · NeurIPS 2023 9 · Sungwon Kim ~Sungwon_Kim2, Kevin J. Shih, Rohan Badlani, Joao Felipe Santos, Evelina Bakhturina, Mikyas T. Desta, Rafael Valle, Sungroh Yoon, Bryan Catanzaro

While recent large-scale neural codec language models have shown significant improvement in zero-shot TTS by training on thousands of hours of data, they suffer from drawbacks such as a lack of robustness, slow sampling speed similar to previous autoregressive TTS methods, and reliance on pre-trained neural codec representations. Our work proposes P-Flow, a fast and data-efficient zero-shot TTS model that uses speech prompts for speaker adaptation. P-Flow comprises a speech-prompted text encoder for speaker adaptation and a flow matching generative decoder for high-quality and fast speech synthesis. Our speech-prompted text encoder uses speech prompts and text input to generate speaker-conditional text representation. The flow matching generative decoder uses the speaker-conditional output to synthesize high-quality personalized speech significantly faster than in real-time. Unlike the neural codec language models, we specifically train P-Flow on LibriTTS dataset using a continuous mel-representation. Through our training method using continuous speech prompts, P-Flow matches the speaker similarity performance of the large-scale zero-shot TTS models with two orders of magnitude less training data and has more than 20 faster sampling speed. Our results show that P-Flow has better pronunciation and is preferred in human likeness and speaker similarity to its recent state-of-the-art counterparts, thus defining P-Flow as an attractive and desirable alternative.

📄 PDF Abstract BibTeX

Code (1)

p0p4k/pflowtts_pytorch pytorch

Tasks

DecoderSpeech Synthesis

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

2025-06-16 · Han Zhu, Wei Kang, Zengwei Yao, Liyong Guo 외

Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-qualit…

DecoderSpeech Synthesistext-to-speechText to Speech+1

COT Flow: Learning Optimal-Transport Image Sampling and Editing by Contrastive Pairs

2024-06-17 · Xinrui Zu, Qian Tao

Diffusion models have demonstrated strong performance in sampling and editing multi-modal data with high generation quality, yet they suffer from the iterative generation process which is computationally expensive and sl…

Translation

BLISSNet: Deep Operator Learning for Fast and Accurate Flow Reconstruction from Sparse Sensor Measurements

2026-02-27 · Maksym Veremchuk, K. Andrea Scott, Zhao Pan arxiv

Reconstructing fluid flows from sparse sensor measurements is a fundamental challenge in science and engineering. Widely separated measurements and complex, multiscale dynamics make accurate recovery of fine-scale struct…

Zero-shot GeneralizationComputational Efficiency

StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

2024-12-06 · Jixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning 외

Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using langu…

Voice Conversion

From Flows to Words: Can Zero-/Few-Shot LLMs Detect Network Intrusions? A Grammar-Constrained, Calibrated Evaluation on UNSW-NB15

2025-10-18 · Mohammad Abdul Rehman, Syed Imad Ali Shah, Abbas Anwar, Noor Islam arxiv

Large Language Models (LLMs) can reason over natural-language inputs, but their role in intrusion detection without fine-tuning remains uncertain. This study evaluates a prompt-only approach on UNSW-NB15 by converting ea…

Intrusion Detection