paper-with-me

홈 › Papers

Audio ControlNet for Fine-Grained Audio Generation and Editing

2026-02-04 · Haina Zhu, Yao Xiao, Xiquan Li, Ziyang Ma, Jianwei Yu, Bowen Zhang, Mingqi Yang, Xie Chen arxiv

We study the fine-grained text-to-audio (T2A) generation task. While recent models can synthesize high-quality audio from text descriptions, they often lack precise control over attributes such as loudness, pitch, and sound events. Unlike prior approaches that retrain models for specific control types, we propose to train ControlNet models on top of pre-trained T2A backbones to achieve controllable generation over loudness, pitch, and event roll. We introduce two designs, T2A-ControlNet and T2A-Adapter, and show that the T2A-Adapter model offers a more efficient structure with strong control ability. With only 38M additional parameters, T2A-Adapter achieves state-of-the-art performance on the AudioSet-Strong in both event-level and segment-level F1 scores. We further extend this framework to audio editing, proposing T2A-Editor for removing and inserting audio events at time locations specified by instructions. Models, code, dataset pipelines, and benchmarks will be released to support future research on controllable audio generation and editing.

📄 PDF Abstract BibTeX arXiv:2602.04680

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Similar Papers 제목 키워드 기반

LiLAC: A Lightweight Latent ControlNet for Musical Audio Generation

2025-06-13 · Tom Baker, Javier Nistal

Text-to-audio diffusion models produce high-quality and diverse music but many, if not most, of the SOTA models lack the fine-grained, time-varying controls essential for music production. ControlNet enables attaching ex…

Audio Generation

Music ControlNet: A model similar to SD ControlNetD that can accurately control music generation

2023-11-07 · . 2023 11 · Wu, Shih-Lun and Donahue, Chris and Watanabe, Shinji and Bryan 외

Text-to-music generation models are now capable of generating high-quality music audio in broad styles. However, text control is primarily suitable for the manipulation of global musical attributes like genre, mood, and …

Music GenerationRhythmText-to-Music Generation

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

2025-05-22 · Zhi Zhong, Akira Takahashi, Shuyang Cui, Keisuke Toyama 외

Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the task has gained increasing attention in th…

Audio Synthesis

Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer

2024-10-07 · Siyuan Hou, Shansong Liu, Ruibin Yuan, Wei Xue 외

Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structu…

Music GenerationMusic Style TransferStyle TransferText-to-Music Generation

MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners

2025-06-23 · Fang-Duo Tsai, Shih-Lun Wu, Weijaw Lee, Sheng-Ping Yang 외

We propose MuseControlLite, a lightweight mechanism designed to fine-tune text-to-music generation models for precise conditioning using various time-varying musical attributes and reference audio signals. The key findin…

AttributeAudio inpaintingMusic GenerationText-to-Music Generation