paper-with-me

홈 › Papers

Translate your gibberish: black-box adversarial attack on machine translation systems

2023-03-20 · Andrei Chertkov, Olga Tsymboi, Mikhail Pautov, Ivan Oseledets

Neural networks are deployed widely in natural language processing tasks on the industrial scale, and perhaps the most often they are used as compounds of automatic machine translation systems. In this work, we present a simple approach to fool state-of-the-art machine translation tools in the task of translation from Russian to English and vice versa. Using a novel black-box gradient-free tensor-based optimizer, we show that many online translation tools, such as Google, DeepL, and Yandex, may both produce wrong or offensive translations for nonsensical adversarial input queries and refuse to translate seemingly benign input phrases. This vulnerability may interfere with understanding a new language and simply worsen the user's experience while using machine translation systems, and, hence, additional improvements of these tools are required to establish better translation.

📄 PDF Abstract BibTeX arXiv:2303.10974

Code (1)

andreichertkov/tranfighterpro 공식 구현 pytorch

Tasks

Adversarial AttackMachine TranslationTranslation

Similar Papers 제목 키워드 기반

AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts

2024-10-29 · Vishal Kumar, Zeyi Liao, Jaylen Jones, Huan Sun

Although large language models (LLMs) are typically aligned, they remain vulnerable to jailbreaking through either carefully crafted prompts in natural language or, interestingly, gibberish adversarial suffixes. However,…

Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models

2026-05-25 · Arian Komaei Koma, Seyed Amir Kasaei, AmirMahdi Sadeghzadeh, Mohammad Hossein Rohban arxiv

Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been introduced to make the model generate such unlearned concepts. These …

AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

2023-10-23 · Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu 외

Safety alignment of Large Language Models (LLMs) can be compromised with manual jailbreak attacks and (automatic) adversarial attacks. Recent studies suggest that defending against these attacks is possible: adversarial …

Adversarial AttackBlockingSafety Alignment

Boosting the Transferability of Video Adversarial Examples via Temporal Translation

2021-10-18 · Zhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang

Although deep-learning based video recognition models have achieved remarkable success, they are vulnerable to adversarial examples that are generated by adding human-imperceptible perturbations on clean video samples. A…

Adversarial AttackTranslationVideo Recognition

Are Deep Speech Denoising Models Robust to Adversarial Noise?

2025-03-14 · Will Schwarzer, Philip S. Thomas, Andrea Fanelli, Xiaoyu Liu

Deep noise suppression (DNS) models enjoy widespread use throughout a variety of high-stakes speech applications. However, in this paper, we show that four recent DNS models can each be reduced to outputting unintelligib…

DenoisingSpeech Denoising