paper-with-me

홈 › Papers

Transformation of audio embeddings into interpretable, concept-based representations

2025-04-18 · Alice Zhang, Edison Thomaz, Lie Lu

Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to interpret the information encoded in their internal audio representations. In this work, we explore the semantic interpretability of audio embeddings extracted from these neural networks by leveraging CLAP, a contrastive learning model that brings audio and text into a shared embedding space. We implement a post-hoc method to transform CLAP embeddings into concept-based, sparse representations with semantic interpretability. Qualitative and quantitative evaluations show that the concept-based representations outperform or match the performance of original audio embeddings on downstream tasks while providing interpretability. Additionally, we demonstrate that fine-tuning the concept-based representations can further improve their performance on downstream tasks. Lastly, we publish three audio-specific vocabularies for concept-based interpretability of audio embeddings.

📄 PDF Abstract BibTeX arXiv:2504.14076

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

2026-05-28 · Yonggang Zhu, Liting Gao, Aidong Men, Wenwu Wang arxiv

Contrastive Language-Audio Pretraining (CLAP) models are widely used for audio understanding and support modality-agnostic condition swapping in many zero-shot applications. However, their performance is heavily affected…

Zero-shot Audio CaptioningDimensionality Reduction

Finding Interpretable Concept Spaces in Node Embeddings using Knowledge Bases

2019-10-11 · Maximilian Idahl, Megha Khosla, Avishek Anand

In this paper we propose and study the novel problem of explaining node embeddings by finding embedded human interpretable subspaces in already trained unsupervised node representation embeddings. We use an external know…

Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models

2026-05-21 · Piotr Kubaty, Patryk Marszałek, Łukasz Struski, Adam Wróbel 외 arxiv

Vision-language models learn powerful multimodal embeddings, yet their internal semantics remain opaque. While sparse autoencoders (SAEs) can extract interpretable features, they rely on expanding the representation dime…

Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects

2024-09-27 · Annie Chu, Patrick O'Reilly, Julia Barnett, Bryan Pardo

This work introduces Text2FX, a method that leverages CLAP embeddings and differentiable digital signal processing to control audio effects, such as equalization and reverberation, using open-vocabulary natural language …

Are Nearby Neighbors Relatives?: Testing Deep Music Embeddings

2019-04-15 · Jaehun Kim, Julián Urbano, Cynthia C. S. Liem, Alan Hanjalic

Deep neural networks have frequently been used to directly learn representations useful for a given task from raw input data. In terms of overall performance metrics, machine learning solutions employing deep representat…