paper-with-me

Papers

Audio-Visual Event Recognition through the lens of Adversary

2020-11-15 · Juncheng B Li, Kaixin Ma, Shuhui Qu, Po-Yao Huang, Florian Metze

As audio/visual classification models are widely deployed for sensitive tasks like content filtering at scale, it is critical to understand their robustness along with improving the accuracy. This work aims to study several key questions related to multimodal learning through the lens of adversarial noises: 1) The trade-off between early/middle/late fusion affecting its robustness and accuracy 2) How do different frequency/time domain features contribute to the robustness? 3) How do different neural modules contribute to the adversarial noise? In our experiment, we construct adversarial examples to attack state-of-the-art neural models trained on Google AudioSet. We compare how much attack potency in terms of adversarial perturbation of size $\epsilon$ using different $L_p$ norms we would need to "deactivate" the victim model. Using adversarial noise to ablate multimodal models, we are able to provide insights into what is the best potential fusion strategy to balance the model parameters/accuracy and robustness trade-off and distinguish the robust features versus the non-robust features that various neural networks model tend to learn.

📄 PDF Abstract BibTeX arXiv:2011.07430

Code (1)

lijuncheng16/AudioSetDoneRight 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Multi-level Attention Fusion Network for Audio-visual Event Recognition

2021-06-12 · Mathilde Brousmiche, Jean Rouat, Stéphane Dupont

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level…

Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition

2020-05-18 · ECCV 2020 8 · Di Hu, Xuhong LI, Lichao Mou, Pu Jin 외

Aerial scene recognition is a fundamental task in remote sensing and has recently received increased interest. While the visual information from overhead images with powerful models and efficient algorithms yields consid…

Scene Recognition

Audiovisual Moments in Time: A Large-Scale Annotated Dataset of Audiovisual Actions

2023-08-18 · PLOS ONE 2024 4 · Michael Joannou, Pia Rotshtein, Uta Noppeney

We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Tim…

AENet: Learning Deep Audio Features for Video Analysis

2017-01-03 · Naoya Takahashi, Michael Gygli, Luc van Gool

We propose a new deep network for audio event recognition, called AENet. In contrast to speech, sounds coming from audio events may be produced by a wide variety of sources. Furthermore, distinguishing them often require…

Action RecognitionData AugmentationEvent DetectionGPU+3

Where and When: Space-Time Attention for Audio-Visual Explanations

2021-05-04 · Yanbei Chen, Thomas Hummel, A. Sophia Koepke, Zeynep Akata

Explaining the decision of a multi-modal decision-maker requires to determine the evidence from both modalities. Recent advances in XAI provide explanations for models trained on still images. However, when it comes to m…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)Multimodal Deep Learning