paper-with-me

Papers

TADA! Tuning Audio Diffusion Models through Activation Steering

2026-02-12 · Łukasz Staniszewski, Katarzyna Zaleska, Mateusz Modrzejewski, Kamil Deja arxiv

Audio diffusion models can synthesize high-fidelity music from text, yet achieving fine-grained control over specific musical attributes remains challenging, as their internal mechanisms for representing high-level concepts are poorly understood. In this work, we use activation patching to demonstrate that recent audio diffusion architectures exhibit a semantic bottleneck, where a small, shared subset of consecutive attention layers controls distinct musical concepts, such as the presence of specific instruments, vocals, or genres. Building on this, we systematically evaluate a broad spectrum of steering paradigms, comparing activation steering against prompt-level, score-space, and weight-space interventions, analyzing the interaction between the steering mechanism and the intervention site. Our new benchmark, supported by an extensive user study, demonstrates that localized activation steering establishes a new state-of-the-art in audio concept modulation.

📄 PDF Abstract BibTeX arXiv:2602.11910

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

2024-12-19 · Moayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov 외

We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioni…

Video GenerationVideo Synchronization

PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models

2024-09-20 · Jayneel Vora, Aditya Krishnan, Nader Bouacida, Prabhu RV Shankar 외

Denoising diffusion models have emerged as state-of-the-art in generative tasks across image, audio, and video domains, producing high-quality, diverse, and contextually relevant data. However, their broader adoption is …

Audio GenerationAudio SynthesisDenoisingQuantization

Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning

2024-12-30 · Zixiang Wan, Ziyue Qiu, Yiyang Liu, Wei-Qiang Zhang

Speech Emotion Recognition (SER) involves analyzing vocal expressions to determine the emotional state of speakers, where the comprehensive and thorough utilization of audio information is paramount. Therefore, we propos…

Emotion RecognitionMulti-Task LearningSelf-Supervised LearningSpeech Emotion Recognition

Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval

2024-06-22 · Paul Primus, Gerhard Widmer

Matching raw audio signals with textual descriptions requires understanding the audio's content and the description's semantics and then drawing connections between the two modalities. This paper investigates a hybrid re…

AudioCapsRetrieval

EDIT: Early Diffusion Inference Termination for dLLMs Based on Dynamics of Training Gradients

2025-11-29 · He-Yen Hsieh, Hong Wang, H. T. Kung arxiv

Diffusion-based large language models (dLLMs) refine token generations through iterative denoising, but answers often stabilize before all steps complete. We propose EDIT (Early Diffusion Inference Termination), an infer…