paper-with-me

Papers

AudioGenX: Explainability on Text-to-Audio Generative Models

2025-02-01 · Hyunju Kang, Geonhee Han, Yoonjae Jeong, Hogun Park

Text-to-audio generation models (TAG) have achieved significant advances in generating audio conditioned on text descriptions. However, a critical challenge lies in the lack of transparency regarding how each textual input impacts the generated audio. To address this issue, we introduce AudioGenX, an Explainable AI (XAI) method that provides explanations for text-to-audio generation models by highlighting the importance of input tokens. AudioGenX optimizes an Explainer by leveraging factual and counterfactual objective functions to provide faithful explanations at the audio token level. This method offers a detailed and comprehensive understanding of the relationship between text inputs and audio outputs, enhancing both the explainability and trustworthiness of TAG models. Extensive experiments demonstrate the effectiveness of AudioGenX in producing faithful explanations, benchmarked against existing methods using novel evaluation metrics specifically designed for audio generation tasks.

📄 PDF Abstract BibTeX arXiv:2502.00459

Code (1)

hjkng/audiogenX 공식 구현 pytorch

Tasks

Audio GenerationcounterfactualTAG

Similar Papers 제목 키워드 기반

Explainability Paths for Sustained Artistic Practice with AI

2024-07-21 · Austin Tecks, Thomas Peschlow, Gabriel Vigliensoni

The development of AI-driven generative audio mirrors broader AI trends, often prioritizing immediate accessibility at the expense of explainability. Consequently, integrating such tools into sustained artistic practice …

Play Me Something Icy: Practical Challenges, Explainability and the Semantic Gap in Generative AI Music

2024-08-13 · Jesse Allison, Drew Farrar, Treya Nash, Carlos Román 외

This pictorial aims to critically consider the nature of text-to-audio and text-to-music generative tools in the context of explainable AI. As a group of experimental musicians and researchers, we are enthusiastic about …

mllm-shap: A Shapley Value Explainability Platform for Text-Audio Multimodal Large Language Models

2026-04-21 · Jakub Muszyński, Paweł Pozorski, Maria Ganzha arxiv

We introduce mllm-shap, an open-source Python framework designed to extend Shapley Value (SV) explainability from text-only Large Language Models to Multimodal LLMs (MLLMs) processing joint text and audio inputs. While t…

Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap

2024-10-09 · Georgia Channing, Juil Sock, Ronald Clark, Philip Torr 외

The rapid proliferation of AI-manipulated or generated audio deepfakes poses serious challenges to media integrity and election security. Current AI-driven detection solutions lack explainability and underperform in real…

Audio Deepfake DetectionDeepFake DetectionFace Swapping

A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations

2025-06-03 · Petr Grinberg, Ankur Kumar, Surya Koppisetti, Gaurav Bharaj

Evaluating explainability techniques, such as SHAP and LRP, in the context of audio deepfake detection is challenging due to lack of clear ground truth annotations. In the cases when we are able to obtain the ground trut…

Audio Deepfake DetectionDeepFake DetectionFace Swapping