paper-with-me

Papers

Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing

2024-06-19 · Martin Lebourdais, Théo Mariotte, Antonio Almudévar, Marie Tahon, Alfonso Ortega

Audio segmentation is a key task for many speech technologies, most of which are based on neural networks, usually considered as black boxes, with high-level performances. However, in many domains, among which health or forensics, there is not only a need for good performance but also for explanations about the output decision. Explanations derived directly from latent representations need to satisfy "good" properties, such as informativeness, compactness, or modularity, to be interpretable. In this article, we propose an explainable-by-design audio segmentation model based on non-negative matrix factorization (NMF) which is a good candidate for the design of interpretable representations. This paper shows that our model reaches good segmentation performances, and presents deep analyses of the latent representation extracted from the non-negative matrix. The proposed approach opens new perspectives toward the evaluation of interpretable representations according to "good" properties.

📄 PDF Abstract BibTeX arXiv:2406.13385

Code (1)

Lebourdais/3MAS 공식 구현 pytorch

Tasks

InformativenessSegmentation

Similar Papers 제목 키워드 기반

An Explainable Proxy Model for Multiabel Audio Segmentation

2024-01-16 · Théo Mariotte, Antonio Almudévar, Marie Tahon, Alfonso Ortega

Audio signal segmentation is a key task for automatic audio indexing. It consists of detecting the boundaries of class-homogeneous segments in the signal. In many applications, explainable AI is a vital process for trans…

Decision MakingSegmentation

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

2025-02-01 · Jia Li, Wenjie Zhao, Ziru Huang, Yunhui Guo 외

Unlike traditional visual segmentation, audio-visual segmentation (AVS) requires the model not only to identify and segment objects but also to determine whether they are sound sources. Recent AVS approaches, leveraging …

Segmentation

Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment

2025-03-17 · CVPR 2025 1 · Chen Liu, Peike Li, Liying Yang, Dadong Wang 외

Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from …

Contrastive Learning

Only Positive Cases: 5-fold High-order Attention Interaction Model for Skin Segmentation Derived Classification

2023-11-27 · Renkai Wu, Yinghao Liu, Pengchen Liang, Qing Chang

Computer-aided diagnosis of skin diseases is an important tool. However, the interpretability of computer-aided diagnosis is currently poor. Dermatologists and patients cannot intuitively understand the learning and pred…

Image SegmentationLesion ClassificationLesion SegmentationMedical Image Classification+3

Complex-valued neural networks for voice anti-spoofing

2023-08-22 · Nicolas M. Müller, Philip Sperl, Konstantin Böttinger

Current anti-spoofing and audio deepfake detection systems use either magnitude spectrogram-based features (such as CQT or Melspectrograms) or raw audio processed through convolution or sinc-layers. Both methods have dra…

Audio Deepfake DetectionDeepFake DetectionFace SwappingVoice Anti-spoofing