paper-with-me

홈 › Papers

Audio Explanation Synthesis with Generative Foundation Models

2024-10-10 · Alican Akman, Qiyang Sun, Björn W. Schuller

The increasing success of audio foundation models across various tasks has led to a growing need for improved interpretability to understand their intricate decision-making processes better. Existing methods primarily focus on explaining these models by attributing importance to elements within the input space based on their influence on the final decision. In this paper, we introduce a novel audio explanation method that capitalises on the generative capacity of audio foundation models. Our method leverages the intrinsic representational power of the embedding space within these models by integrating established feature attribution techniques to identify significant features in this space. The method then generates listenable audio explanations by prioritising the most important features. Through rigorous benchmarking against standard datasets, including keyword spotting and speech emotion recognition, our model demonstrates its efficacy in producing audio explanations.

📄 PDF Abstract BibTeX arXiv:2410.07530

Code (1)

glam-imperial/AudioXgen 공식 구현 pytorch

Tasks

BenchmarkingDecision MakingEmotion RecognitionKeyword SpottingSpeech Emotion Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond

2025-05-07 · Jessie Richter-Powell, Antonio Torralba, Jonathan Lorraine

We introduce Audio-SDS, a generalization of Score Distillation Sampling (SDS) to text-conditioned audio diffusion models. While SDS was initially designed for text-to-3D generation using image diffusion, its core idea of…

3D GenerationAudio Source SeparationText to 3D

TRACE: Training-Free Partial Audio Deepfake Detection via Embedding Trajectory Analysis of Speech Foundation Models

2026-04-01 · Awais Khan, Muhammad Umar Farooq, Kutub Uddin, Khalid Malik arxiv

Partial audio deepfakes, where synthesized segments are spliced into genuine recordings, are particularly deceptive because most of the audio remains authentic. Existing detectors are supervised: they require frame-level…

Audio Deepfake Detection

The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models

2025-09-30 · Lina Conti, Dennis Fucci, Marco Gaido, Matteo Negri 외 arxiv

Contrastive explanations, which indicate why an AI system produced one output (the target) instead of another (the foil), are widely regarded in explainable AI as more informative and interpretable than standard explanat…

LMAC-TD: Producing Time Domain Explanations for Audio Classifiers

2024-09-13 · Eleonora Mancini, Francesco Paissan, Mirco Ravanelli, Cem Subakan

Neural networks are typically black-boxes that remain opaque with regards to their decision mechanisms. Several works in the literature have proposed post-hoc explanation methods to alleviate this issue. This paper propo…

Decoder

AudioLCM: Text-to-Audio Generation with Latent Consistency Models

2024-06-01 · Huadai Liu, Rongjie Huang, Yang Liu, Hengyuan Cao 외

Recent advancements in Latent Diffusion Models (LDMs) have propelled them to the forefront of various generative tasks. However, their iterative sampling process poses a significant computational burden, resulting in slo…

Audio GenerationAudio SynthesisGPU