paper-with-me

Papers

Low-Resource Guidance for Controllable Latent Audio Diffusion

2026-03-04 · Zachary Novack, Zack Zukowski, CJ Carr, Julian Parker, Zach Evans, Josiah Taylor, Taylor Berg-Kirkpatrick, Julian McAuley, Jordi Pons arxiv

Generative audio requires fine-grained controllable outputs, yet most existing methods require model retraining on specific controls or inference-time controls (\textit{e.g.}, guidance) that can also be computationally demanding. By examining the bottlenecks of existing guidance-based controls, in particular their high cost-per-step due to decoder backpropagation, we introduce a guidance-based approach through selective TFG and Latent-Control Heads (LatCHs), which enables controlling latent audio diffusion models with low computational overhead. LatCHs operate directly in latent space, avoiding the expensive decoder step, and requiring minimal training resources (7M parameters and $\approx$ 4 hours of training). Experiments with Stable Audio Open demonstrate effective control over intensity, pitch, and beats (and a combination of those) while maintaining generation quality. Our method balances precision and audio fidelity with far lower computational costs than standard end-to-end guidance. Demo examples can be found at https://zacharynovack.github.io/latch/latch.html.

📄 PDF Abstract BibTeX arXiv:2603.04366

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bass Accompaniment Generation via Latent Diffusion

2024-02-02 · Marco Pasini, Maarten Grachten, Stefan Lattner

The ability to automatically generate music that appropriately matches an arbitrary input track is a challenging task. We present a novel controllable system for generating single stems to accompany musical mixes of arbi…

Audio Generation

Controllable Music Production with Diffusion Models and Guidance Gradients

2023-11-01 · Mark Levy, Bruno Di Giorgi, Floris Weers, Angelos Katharopoulos 외

We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider…

Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation

2024-10-14 · Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye 외

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. C…

Audio Generationmultimodal generation

Example-Based Framework for Perceptually Guided Audio Texture Generation

2023-08-23 · Purnima Kamath, Chitralekha Gupta, Lonce Wyse, Suranga Nanayakkara

Controllable generation using StyleGANs is usually achieved by training the model using labeled data. For audio textures, however, there is currently a lack of large semantically labeled datasets. Therefore, to control g…

AttributeTexture Synthesis

Controllable Stylistic Text Generation with Train-Time Attribute-Regularized Diffusion

2025-10-07 · Fan Zhou, Chang Tian, Tim Van de Cruys arxiv

Generating stylistic text with specific attributes is a key problem in controllable text generation. Recently, diffusion models have emerged as a powerful paradigm for both visual and textual generation. Existing approac…

Text Generation