AudioFool: Fast, Universal and synchronization-free Cross-Domain Attack on Speech Recognition
Automatic Speech Recognition systems have been shown to be vulnerable to adversarial attacks that manipulate the command executed on the device. Recent research has focused on exploring methods to create such attacks, however, some issues relating to Over-The-Air (OTA) attacks have not been properly addressed. In our work, we examine the needed properties of robust attacks compatible with the OTA model, and we design a method of generating attacks with arbitrary such desired properties, namely the invariance to synchronization, and the robustness to filtering: this allows a Denial-of-Service (DoS) attack against ASR systems. We achieve these characteristics by constructing attacks in a modified frequency domain through an inverse Fourier transform. We evaluate our method on standard keyword classification tasks and analyze it in OTA, and we analyze the properties of the cross-domain attacks to explain the efficiency of the approach.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
The classical mean negative asynchrony in sensorimotor synchronization is not universal in humans. A cross-cultural study
The present study examines to what extent cultural background determines sensorimotor synchronization in humans
Fast Sparsely Synchronized Brain Rhythms in A Scale-Free Neural Network
We consider a directed Barab\'{a}si-Albert scale-free network model with symmetric preferential attachment with the same in- and out-degrees, and study emergence of sparsely synchronized rhythms for a fixed attachment de…
RhythmScale-free Protocol Design for Output Synchronization of Heterogeneous Multi-agent subject to Unknown, Non-uniform and Arbitrarily Large Input Delays
This paper studies output synchronization problems for heterogeneous networks of continuous- or discrete-time right-invertible linear agents in presence of unknown, non-uniform and arbitrarily large input delay based on …
HDTR-Net: A Real-Time High-Definition Teeth Restoration Network for Arbitrary Talking Face Generation Methods
Talking Face Generation (TFG) aims to reconstruct facial movements to achieve high natural lip movements from audio and facial features that are under potential connections. Existing TFG methods have made significant adv…
Face GenerationSuper-ResolutionTalking Face GenerationUniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control
Video Diffusion Models have been developed for video generation, usually integrating text and image conditioning to enhance control over the generated content. Despite the progress, ensuring consistency across frames rem…
DiversityVideo Generation