Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
Audio segmentation is a key task for many speech technologies, most of which are based on neural networks, usually considered as black boxes, with high-level performances. However, in many domains, among which health or forensics, there is not only a need for good performance but also for explanations about the output decision. Explanations derived directly from latent representations need to satisfy "good" properties, such as informativeness, compactness, or modularity, to be interpretable. In this article, we propose an explainable-by-design audio segmentation model based on non-negative matrix factorization (NMF) which is a good candidate for the design of interpretable representations. This paper shows that our model reaches good segmentation performances, and presents deep analyses of the latent representation extracted from the non-negative matrix. The proposed approach opens new perspectives toward the evaluation of interpretable representations according to "good" properties.
Code (1)
Tasks
InformativenessSegmentationSimilar Papers 제목 키워드 기반
An Explainable Proxy Model for Multiabel Audio Segmentation
Audio signal segmentation is a key task for automatic audio indexing. It consists of detecting the boundaries of class-homogeneous segments in the signal. In many applications, explainable AI is a vital process for trans…
Decision MakingSegmentationDo Audio-Visual Segmentation Models Truly Segment Sounding Objects?
Unlike traditional visual segmentation, audio-visual segmentation (AVS) requires the model not only to identify and segment objects but also to determine whether they are sound sources. Recent AVS approaches, leveraging …
SegmentationRobust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from …
Contrastive LearningOnly Positive Cases: 5-fold High-order Attention Interaction Model for Skin Segmentation Derived Classification
Computer-aided diagnosis of skin diseases is an important tool. However, the interpretability of computer-aided diagnosis is currently poor. Dermatologists and patients cannot intuitively understand the learning and pred…
Image SegmentationLesion ClassificationLesion SegmentationMedical Image Classification+3Complex-valued neural networks for voice anti-spoofing
Current anti-spoofing and audio deepfake detection systems use either magnitude spectrogram-based features (such as CQT or Melspectrograms) or raw audio processed through convolution or sinc-layers. Both methods have dra…
Audio Deepfake DetectionDeepFake DetectionFace SwappingVoice Anti-spoofing