paper-with-me

Papers

A Study on Robustness to Perturbations for Representations of Environmental Sound

2022-03-20 · Sangeeta Srivastava, Ho-Hsiang Wu, Joao Rulff, Magdalena Fuentes, Mark Cartwright, Claudio Silva, Anish Arora, Juan Pablo Bello

Audio applications involving environmental sound analysis increasingly use general-purpose audio representations, also known as embeddings, for transfer learning. Recently, Holistic Evaluation of Audio Representations (HEAR) evaluated twenty-nine embedding models on nineteen diverse tasks. However, the evaluation's effectiveness depends on the variation already captured within a given dataset. Therefore, for a given data domain, it is unclear how the representations would be affected by the variations caused by myriad microphones' range and acoustic conditions -- commonly known as channel effects. We aim to extend HEAR to evaluate invariance to channel effects in this work. To accomplish this, we imitate channel effects by injecting perturbations to the audio signal and measure the shift in the new (perturbed) embeddings with three distance measures, making the evaluation domain-dependent but not task-dependent. Combined with the downstream performance, it helps us make a more informed prediction of how robust the embeddings are to the channel effects. We evaluate two embeddings -- YAMNet, and OpenL3 on monophonic (UrbanSound8K) and polyphonic (SONYC-UST) urban datasets. We show that one distance measure does not suffice in such task-independent evaluation. Although Fr\'echet Audio Distance (FAD) correlates with the trend of the performance drop in the downstream task most accurately, we show that we need to study FAD in conjunction with the other distances to get a clear understanding of the overall effect of the perturbation. In terms of the embedding performance, we find OpenL3 to be more robust than YAMNet, which aligns with the HEAR evaluation.

📄 PDF Abstract BibTeX arXiv:2203.10425

Code (0)

등록된 구현이 없습니다.

Tasks

FADTransfer Learning

Similar Papers 제목 키워드 기반

From Sound Representation to Model Robustness

2020-07-27 · Mohammad Esmaeilpour, Patrick Cardinal, Alessandro Lameiras Koerich

In this paper, we investigate the impact of different standard environmental sound representations (spectrograms) on the recognition performance and adversarial attack robustness of a victim residual convolutional neural…

Adversarial AttackAdversarial RobustnessBenchmarkingmodel

BEAT2AASIST model with layer fusion for ESDD 2026 Challenge

2025-12-17 · Sanghyeok Chung, Eujin Kim, Donggun Kim, Gaeun Heo 외 arxiv

Recent advances in audio generation have increased the risk of realistic environmental sound manipulation, motivating the ESDD 2026 Challenge as the first large-scale benchmark for Environmental Sound Deepfake Detection …

DeepFake DetectionData AugmentationAudio Generation

Environmental Sound Classification with Parallel Temporal-spectral Attention

2019-12-14 · Helin Wang, Yuexian Zou, Dading Chong, Wenwu Wang

Convolutional neural networks (CNN) are one of the best-performing neural network architectures for environmental sound classification (ESC). Recently, temporal attention mechanisms have been used in CNN to capture the u…

Acoustic Scene ClassificationAudio ClassificationClassificationEnvironmental Sound Classification+3

From Environmental Sound Representation to Robustness of 2D CNN Models Against Adversarial Attacks

2022-04-14 · Mohammad Esmaeilpour, Patrick Cardinal, Alessandro Lameiras Koerich

This paper investigates the impact of different standard environmental sound representations (spectrograms) on the recognition performance and adversarial attack robustness of a victim residual convolutional neural netwo…

Adversarial AttackAdversarial RobustnessBenchmarking

Towards Robust Deep Reinforcement Learning against Environmental State Perturbation

2025-06-10 · Chenxu Wang, Huaping Liu

Adversarial attacks and robustness in Deep Reinforcement Learning (DRL) have been widely studied in various threat models; however, few consider environmental state perturbations, which are natural in embodied scenarios.…

Deep Reinforcement Learningreinforcement-learningReinforcement Learning