paper-with-me

Papers

Push-Pull: Characterizing the Adversarial Robustness for Audio-Visual Active Speaker Detection

2022-10-03 · Xuanjun Chen, Haibin Wu, Helen Meng, Hung-Yi Lee, Jyh-Shing Roger Jang

Audio-visual active speaker detection (AVASD) is well-developed, and now is an indispensable front-end for several multi-modal applications. However, to the best of our knowledge, the adversarial robustness of AVASD models hasn't been investigated, not to mention the effective defense against such attacks. In this paper, we are the first to reveal the vulnerability of AVASD models under audio-only, visual-only, and audio-visual adversarial attacks through extensive experiments. What's more, we also propose a novel audio-visual interaction loss (AVIL) for making attackers difficult to find feasible adversarial examples under an allocated attack budget. The loss aims at pushing the inter-class embeddings to be dispersed, namely non-speech and speech clusters, sufficiently disentangled, and pulling the intra-class embeddings as close as possible to keep them compact. Experimental results show the AVIL outperforms the adversarial training by 33.14 mAP (%) under multi-modal attacks.

📄 PDF Abstract BibTeX arXiv:2210.00753

Code (0)

등록된 구현이 없습니다.

Tasks

Active Speaker DetectionAdversarial RobustnessAudio-Visual Active Speaker Detection

Similar Papers 제목 키워드 기반

PushPull-Net: Inhibition-driven ResNet robust to image corruptions

2024-08-07 · Guru Swaroop Bennabhaktula, Enrique Alegre, Nicola Strisciuglio, George Azzopardi

We introduce a novel computational unit, termed PushPull-Conv, in the first layer of a ResNet architecture, inspired by the anti-phase inhibition phenomenon observed in the primary visual cortex. This unit redefines the …

Data AugmentationDomain Generalization

A Push-Pull Layer Improves Robustness of Convolutional Neural Networks

2019-01-29 · Nicola Strisciuglio, Manuel Lopez-Antequera, Nicolai Petkov

We propose a new layer in Convolutional Neural Networks (CNNs) to increase their robustness to several types of noise perturbations of the input images. We call this a push-pull layer and compute its response as the comb…

General Classificationimage-classificationImage Classification

Inhibition-augmented ConvNets

2021-01-01 · Nicola Strisciuglio, George Azzopardi, Nicolai Petkov

Convolutional Networks (ConvNets) suffer from insufficient robustness to common corruptions and perturbations of the input, unseen during training. We address this problem by including a form of response inhibition in…

Characterizing Audio Adversarial Examples Using Temporal Dependency

2018-09-28 · ICLR 2019 5 · Zhuolin Yang, Bo Li, Pin-Yu Chen, Dawn Song

Recent studies have highlighted adversarial examples as a ubiquitous threat to different neural network models and many downstream applications. Nonetheless, as unique data properties have inspired distinct and powerful …

Adversarial DefenseAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1

Brain-inspired robust delineation operator

2018-11-26 · Nicola Strisciuglio, George Azzopardi, Nicolai Petkov

In this paper we present a novel filter, based on the existing COSFIRE filter, for the delineation of patterns of interest. It includes a mechanism of push-pull inhibition that improves robustness to noise in terms of sp…