paper-with-me

홈 › Papers

Turn-Based Structural Triggers: Structure-Conditioned Backdoors in Multi-Turn LLMs

2026-01-20 · Yiyang Lu, Jinwen He, Yue Zhao, Kai Chen, Ruigang Liang, Cheng Hong, Yingjun Zhang arxiv

Large Language Models (LLMs) are increasingly deployed as multi-turn assistants and customized through instruction tuning with project-specific training components. This practice creates a supply-chain risk when organizations reuse third-party fine-tuning frameworks, trainer extensions, or outsourced training code: an adversary who subtly compromises the loss-computation component can inject malicious supervision during fine-tuning while leaving the stored training corpus, model architecture, and deployment interface unchanged. Existing LLM backdoors and defenses are largely prompt-centric, relying on lexical, syntactic, or semantic patterns in user inputs while overlooking structural signals in multi-turn conversations. We propose Turn-based Structural Trigger (TST), a prompt-free backdoor that uses dialogue turn position as its activation condition. TST exploits structural cues implicitly encoded by chat templates and is implanted without modifying the stored dialogue corpus. Its trigger is automatically present once the conversation reaches the attacker-specified turn, making activation independent of downstream user inputs and resistant to prompt filtering, sanitization, and paraphrasing. In our primary setting, the model behaves normally during early interactions and activates the attacker-defined behavior only after the conversation reaches the designated structural condition. Across four open-source LLM families, TST achieves an average Attack Success Rate of 98.10% on target turns and a Clean Rate of 99.96% on non-target turns, while retaining 97.78% of clean-model utility. These results identify dialogue structure as an overlooked attack surface and motivate structure-aware auditing beyond prompt inspection.

📄 PDF Abstract BibTeX arXiv:2601.14340

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mechanistic Exploration of Backdoored Large Language Model Attention Patterns

2025-08-19 · Mohammed Abu Baker, Lakshmi Babu-Saheer arxiv

Backdoor attacks creating 'sleeper agents' in large language models (LLMs) pose significant safety risks. This study employs mechanistic interpretability to explore resulting internal structural differences. Comparing cl…

Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs

2026-06-02 · Lisa Bouger, Théo Lasnier, Philippe Loubet Moundi, Yannick Teglia 외 arxiv

Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses target backdoors one at a time and typically require knowledge of the…

Continual Pretraining

MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs

2026-05-14 · Rui Wen, Mark Russinovich, Andrew Paverd, Jun Sakuma 외 arxiv

Backdoor attacks pose a serious security threat to large language models (LLMs), which are increasingly deployed as general-purpose assistants in safety- and privacy-critical applications. Existing LLM backdoors rely pri…

Manipulating Trajectory Prediction with Backdoors

2023-12-21 · Kaouther Messaoud, Kathrin Grosse, Mickael Chen, Matthieu Cord 외

Autonomous vehicles ought to predict the surrounding agents' trajectories to allow safe maneuvers in uncertain and complex traffic situations. As companies increasingly apply trajectory prediction in the real world, secu…

Autonomous VehiclesPredictionTrajectory Prediction

Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors

2026-03-10 · Gorka Abad, Ermes Franch, Stefanos Koffas, Stjepan Picek arxiv

Current backdoor defenses assume that neutralizing a known trigger removes the backdoor. We show this trigger-centric view is incomplete: \emph{alternative triggers}, patterns perceptually distinct from training triggers…