paper-with-me

Papers

Attack-in-the-Chain: Bootstrapping Large Language Models for Attacks Against Black-box Neural Ranking Models

2024-12-25 · Yu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, Xueqi Cheng

Neural ranking models (NRMs) have been shown to be highly effective in terms of retrieval performance. Unfortunately, they have also displayed a higher degree of sensitivity to attacks than previous generation models. To help expose and address this lack of robustness, we introduce a novel ranking attack framework named Attack-in-the-Chain, which tracks interactions between large language models (LLMs) and NRMs based on chain-of-thought (CoT) prompting to generate adversarial examples under black-box settings. Our approach starts by identifying anchor documents with higher ranking positions than the target document as nodes in the reasoning chain. We then dynamically assign the number of perturbation words to each node and prompt LLMs to execute attacks. Finally, we verify the attack performance of all nodes at each reasoning step and proceed to generate the next reasoning step. Empirical results on two web search benchmarks show the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2412.18770

Code (1)

Davion-Liu/AttChain 공식 구현 pytorch

Similar Papers 제목 키워드 기반

(In)Stability for the Blockchain: Deleveraging Spirals and Stablecoin Attacks

2019-06-05 · Ariah Klages-Mundt, Andreea Minca

We develop a model of stable assets, including non-custodial stablecoins backed by cryptocurrencies. Such stablecoins are popular methods for bootstrapping price stability within public blockchain settings. We derive fun…

Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection

2024-04-07 · Zhilong Wang, Yebo Cao, Peng Liu

Jailbreak attacks on Language Model Models (LLMs) entail crafting prompts aimed at exploiting the models to generate malicious content. Existing jailbreak attacks can successfully deceive the LLMs, however they cannot de…

Language ModelingLanguage Modelling

The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models

2025-02-03 · Zhiyuan Xu, Joseph Gardiner, Sana Belguith

Large language models are typically trained on vast amounts of data during the pre-training phase, which may include some potentially harmful information. Fine-tuning attacks can exploit this by prompting the model to re…

Safety Alignment

Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

2026-06-08 · Blake Bullwinkel, Eugenia Kim, Amanda Minnich, Mark Russinovich arxiv

AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods can produce more robust defenders in tan…

Reinforcement LearningRed Teaming

Reinforcement Learning for Supply Chain Attacks Against Frequency and Voltage Control

2023-09-11 · Amr S. Mohamed, Sumin Lee, Deepa Kundur

The ongoing modernization of the power system, involving new equipment installations and upgrades, exposes the power system to the introduction of malware into its operation through supply chain attacks. Supply chain att…

reinforcement-learning