paper-with-me

Papers

Stable Audio Open

2024-07-19 · Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.

📄 PDF Abstract BibTeX arXiv:2407.14358

Code (1)

stability-ai/stable-audio-tools 공식 구현 pytorch

Tasks

Audio GenerationText-to-Music Generation

Similar Papers 제목 키워드 기반

Woosh: A Sound Effects Foundation Model

2026-04-02 · Gaëtan Hadjeres, Marc Ferras, Khaled Koutini, Benno Weck 외 arxiv

The audio research community depends on open generative models as foundational tools for building novel approaches and establishing baselines. In this report, we present Woosh, Sony AI's publicly released sound effect fo…

MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners

2025-06-23 · Fang-Duo Tsai, Shih-Lun Wu, Weijaw Lee, Sheng-Ping Yang 외

We propose MuseControlLite, a lightweight mechanism designed to fine-tune text-to-music generation models for precise conditioning using various time-varying musical attributes and reference audio signals. The key findin…

AttributeAudio inpaintingMusic GenerationText-to-Music Generation

Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance

2025-09-18 · Francisco Messina, Francesca Ronchini, Luca Comanducci, Paolo Bestagini 외 arxiv

A persistent challenge in generative audio models is data replication, where the model unintentionally generates parts of its training data during inference. In this work, we address this issue in text-to-audio diffusion…

Audio Generation

SAO-Instruct: Free-form Audio Editing using Natural Language Instructions

2025-10-26 · Michael Ungersböck, Florian Grötschla, Luca A. Lanzendörfer, June Young Yi 외 arxiv

Generative models have made significant progress in synthesizing high-fidelity audio from short textual descriptions. However, editing existing audio using natural language has remained largely underexplored. Current app…

ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

2024-01-14 · Yakun Song, Zhuo Chen, Xiaofei Wang, Ziyang Ma 외

The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, existing methods still have some limitation…

Audio GenerationLanguage ModelingLanguage Modellingtext-to-speech+1