paper-with-me

Papers

Listenable Maps for Zero-Shot Audio Classifiers

2024-05-27 · Francesco Paissan, Luca Della Libera, Mirco Ravanelli, Cem Subakan

Interpreting the decisions of deep learning models, including audio classifiers, is crucial for ensuring the transparency and trustworthiness of this technology. In this paper, we introduce LMAC-ZS (Listenable Maps for Audio Classifiers in the Zero-Shot context), which, to the best of our knowledge, is the first decoder-based post-hoc interpretation method for explaining the decisions of zero-shot audio classifiers. The proposed method utilizes a novel loss function that maximizes the faithfulness to the original similarity between a given text-and-audio pair. We provide an extensive evaluation using the Contrastive Language-Audio Pretraining (CLAP) model to showcase that our interpreter remains faithful to the decisions in a zero-shot classification context. Moreover, we qualitatively show that our method produces meaningful explanations that correlate well with different text prompts.

📄 PDF Abstract BibTeX arXiv:2405.17615

Code (0)

등록된 구현이 없습니다.

Tasks

Decoderzero-shot-classificationZero-Shot Learning

Similar Papers 제목 키워드 기반

Listenable Maps for Audio Classifiers

2024-03-19 · Francesco Paissan, Mirco Ravanelli, Cem Subakan

Despite the impressive performance of deep learning models across diverse tasks, their complexity poses challenges for interpretation. This challenge is particularly evident for audio signals, where conveying interpretat…

Decoder

LMAC-TD: Producing Time Domain Explanations for Audio Classifiers

2024-09-13 · Eleonora Mancini, Francesco Paissan, Mirco Ravanelli, Cem Subakan

Neural networks are typically black-boxes that remain opaque with regards to their decision mechanisms. Several works in the literature have proposed post-hoc explanation methods to alleviate this issue. This paper propo…

Decoder

audioLIME: Listenable Explanations Using Source Separation

2020-08-02 · Verena Haunschmid, Ethan Manilow, Gerhard Widmer

Deep neural networks (DNNs) are successfully applied in a wide variety of music information retrieval (MIR) tasks but their predictions are usually not interpretable. We propose audioLIME, a method based on Local Interpr…

Information RetrievalMusic Information RetrievalMusic TaggingRetrieval

MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector

2024-01-10 · Marta R. Costa-jussà, Mariano Coria Meglioli, Pierre Andrews, David Dale 외

Research in toxicity detection in natural language processing for the speech modality (audio-based) is quite limited, particularly for languages other than English. To address these limitations and lay the groundwork for…

Audio Visual Language Maps for Robot Navigation

2023-03-13 · Chenguang Huang, Oier Mees, Andy Zeng, Wolfram Burgard

While interacting in the world is a multi-sensory experience, many robots continue to predominantly rely on visual perception to map and navigate in their environments. In this work, we propose Audio-Visual-Language Maps…

NavigateRobot Navigation