paper-with-me

Papers

Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech

2025-10-03 · Hieu-Nghia Huynh-Nguyen, Huynh Nguyen Dang, Ngoc-Son Nguyen, Van Nguyen arxiv

Zero-shot Text-to-Speech (TTS) has recently advanced significantly, enabling models to synthesize speech from text using short, limited-context prompts. These prompts serve as voice exemplars, allowing the model to mimic speaker identity, prosody, and other traits without extensive speaker-specific data. Although recent approaches incorporating language models, diffusion, and flow matching have proven their effectiveness in zero-shot TTS, they still encounter challenges such as unreliable synthesis caused by token repetition or unexpected content transfer, along with slow inference and substantial computational overhead. Moreover, temporal diversity-crucial for enhancing the naturalness of synthesized speech-remains largely underexplored. To address these challenges, we propose Flamed-TTS, a novel zero-shot TTS framework that emphasizes low computational cost, low latency, and high speech fidelity alongside rich temporal diversity. To achieve this, we reformulate the flow matching training paradigm and incorporate both discrete and continuous representations corresponding to different attributes of speech. Experimental results demonstrate that Flamed-TTS surpasses state-of-the-art models in terms of intelligibility, naturalness, speaker similarity, acoustic characteristics preservation, and dynamic pace. Notably, Flamed-TTS achieves the best WER of 4% compared to the leading zero-shot TTS baselines, while maintaining low latency in inference and high fidelity in generated speech. Code and audio samples are available at our demo page https://flamed-tts.github.io.

📄 PDF Abstract BibTeX arXiv:2510.02848

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer

2022-08-30 · Hongkai Chen, Zixin Luo, Lei Zhou, Yurun Tian 외

Generating robust and reliable correspondences across images is a fundamental task for a diversity of applications. To capture context at both global and local granularity, we propose ASpanFormer, a Transformer-based det…

DiversityHomography Estimation

TFG-Flow: Training-free Guidance in Multimodal Generative Flow

2025-01-24 · Haowei Lin, Shanda Li, Haotian Ye, Yiming Yang 외

Given an unconditional generative model and a predictor for a target property (e.g., a classifier), the goal of training-free guidance is to generate samples with desirable target properties without additional training. …

Drug Design

Dirichlet Flow Matching with Applications to DNA Sequence Design

2024-02-08 · Hannes Stark, Bowen Jing, Chenyu Wang, Gabriele Corso 외

Discrete diffusion or flow models could enable faster and more controllable sequence generation than autoregressive models. We show that na\"ive linear flow matching on the simplex is insufficient toward this goal since …

LSSED: A Robust Segmentation Network for Inflamed Appendix from CT Images

2023-05-05 · ICASSP 2023 5 · Wing W Y. Ng, Peixin Zheng, Ting Wang, Jianjun Zhang 외

Acute appendicitis (AA) is one of the most prevalent surgical acute abdominal condition diseases. The treatment management of A A is highly dependent on the CT image diagnosis. However, the in-flamed appendix exhibits bl…

DecoderManagementSegmentation

SAM-Flow: Source-Anchored Masked Flow for Training-Free Image Editing

2026-06-04 · Haowang Cui, Rui Chen, Tao Luo, Tao Guo 외 arxiv

Training-free image editing has recently attracted increasing attention due to its ability to modify real images using powerful pre-trained diffusion and flow-matching models without additional training. However, existin…

Image Editing