paper-with-me

홈 › Papers

LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation

2024-06-12 · Wenhao Guan, Kaidi Wang, Wangjin Zhou, Yang Wang, Feng Deng, Hui Wang, Lin Li, Qingyang Hong, Yong Qin

Recently, the application of diffusion models has facilitated the significant development of speech and audio generation. Nevertheless, the quality of samples generated by diffusion models still needs improvement. And the effectiveness of the method is accompanied by the extensive number of sampling steps, leading to an extended synthesis time necessary for generating high-quality audio. Previous Text-to-Audio (TTA) methods mostly used diffusion models in the latent space for audio generation. In this paper, we explore the integration of the Flow Matching (FM) model into the audio latent space for audio generation. The FM is an alternative simulation-free method that trains continuous normalization flows (CNF) based on regressing vector fields. We demonstrate that our model significantly enhances the quality of generated audio samples, achieving better performance than prior models. Moreover, it reduces the number of inference steps to ten steps almost without sacrificing performance.

📄 PDF Abstract BibTeX arXiv:2406.08203

Code (1)

gwh22/LAFMA 공식 구현 pytorch

Tasks

Audio Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

PianoKontext: Expressive Performance Rendering from Deadpan Context

2026-06-10 · Dmitrii Gavrilev arxiv

Expressive performance rendering (EPR) aims to generate realistic performances constrained on sequences of notes. However, flow matching audio editing models manipulate only synchronized music samples of the same duratio…

WavFlow: Audio Generation in Waveform Space

2026-05-18 · Feiyan Zhou, Luyuan Wang, Shoufa Chen, Zhe Wang 외 arxiv

Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this work, we challenge this paradigm with WavFlow, a framework that generate…

Audio Generation

DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait Synthesis

2025-10-12 · Peiyin Chen, Zhuowei Yang, Hui Feng, Sheng Jiang 외 arxiv

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-mat…

Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation

2025-06-24 · Jun Wang, Xijuan Zeng, Chunyu Qiang, Ruilong Chen 외

We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce multimodal diffusion transformers to mode…

Audio GenerationAudio-Visual Synchronization

FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait

2024-12-02 · Taekyung Ki, Dongchan Min, Gyeongsu Chae

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling du…

Image AnimationVideo Generation