paper-with-me

Papers

Exploiting Vulnerabilities in Speech Translation Systems through Targeted Adversarial Attacks

2025-03-02 · Chang Liu, Haolin Wu, Xi Yang, Kui Zhang, Cong Wu, Weiming Zhang, Nenghai Yu, Tianwei Zhang, Qing Guo, Jie Zhang

As speech translation (ST) systems become increasingly prevalent, understanding their vulnerabilities is crucial for ensuring robust and reliable communication. However, limited work has explored this issue in depth. This paper explores methods of compromising these systems through imperceptible audio manipulations. Specifically, we present two innovative approaches: (1) the injection of perturbation into source audio, and (2) the generation of adversarial music designed to guide targeted translation, while also conducting more practical over-the-air attacks in the physical world. Our experiments reveal that carefully crafted audio perturbations can mislead translation models to produce targeted, harmful outputs, while adversarial music achieve this goal more covertly, exploiting the natural imperceptibility of music. These attacks prove effective across multiple languages and translation models, highlighting a systemic vulnerability in current ST architectures. The implications of this research extend beyond immediate security concerns, shedding light on the interpretability and robustness of neural speech processing systems. Our findings underscore the need for advanced defense mechanisms and more resilient architectures in the realm of audio systems. More details and samples can be found at https://adv-st.github.io.

📄 PDF Abstract BibTeX arXiv:2503.00957

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

Low-Latency Neural Speech Translation

2018-08-01 · Jan Niehues, Ngoc-Quan Pham, Thanh-Le Ha, Matthias Sperber 외

Through the development of neural machine translation, the quality of machine translation systems has been improved significantly. By exploiting advancements in deep learning, systems are now able to better approximate t…

Machine TranslationMulti-Task LearningNMTSentence+1

Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems

2025-07-21 · Natalia Tomashenko, Emmanuel Vincent, Marc Tommasi arxiv

The temporal dynamics of speech, encompassing variations in rhythm, intonation, and speaking rate, contain important and unique information about speaker identity. This paper proposes a new method for representing speake…

Speaker Verification

Machine Translation of Speech-Like Texts: Strategies for the Inclusion of Context

2017-06-01 · JEPTALNRECITAL 2017 6 · Rachel Bawden

Whilst the focus of Machine Translation (MT) has for a long time been the translation of planned, written texts, more and more research is being dedicated to translating speech-like texts (informal or spontaneous discour…

Machine TranslationTAGTranslation

On Security Weaknesses and Vulnerabilities in Deep Learning Systems

2024-06-12 · Zhongzheng Lai, Huaming Chen, Ruoxi Sun, Yu Zhang 외

The security guarantee of AI-enabled software systems (particularly using deep learning techniques as a functional core) is pivotal against the adversarial attacks exploiting software vulnerabilities. However, little att…

Deep Learning

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

2026-05-29 · Yanjie An, Yuxiang Zhao, Yichi Zhang, Qixi Zheng 외 arxiv

Speech translation systems increasingly span speech-to-text translation (S2TT), speech-to-speech translation (S2ST), offline translation, and streaming generation, producing outputs that differ in modality, speech realiz…

Speech-to-Speech TranslationSpeech-to-Text Translation