paper-with-me

Papers

When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks

2026-02-21 · Shenyang Chen, Liuwan Zhu arxiv

Standard evaluations of backdoor attacks on text-to-image (T2I) models primarily measure trigger activation and visual fidelity. We challenge this paradigm, demonstrating that encoder-side poisoning induces persistent, trigger-free semantic corruption that fundamentally reshapes the representation manifold. We trace this vulnerability to a geometric mechanism: a Jacobian-based analysis reveals that backdoors act as low-rank, target-centered deformations that amplify local sensitivity, causing distortion to propagate coherently across semantic neighborhoods. To rigorously quantify this structural degradation, we introduce SEMAD (Semantic Alignment and Drift), a diagnostic framework that measures both internal embedding drift and downstream functional misalignment. Our findings, validated across diffusion and contrastive paradigms, expose the deep structural risks of encoder poisoning and highlight the necessity of geometric audits beyond simple attack success rates.

📄 PDF Abstract BibTeX arXiv:2602.20193

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs

2024-07-04 · Sara Price, Arjun Panickssery, Sam Bowman, Asa Cooper Stickland

Backdoors are hidden behaviors that are only triggered once an AI system has been deployed. Bad actors looking to create successful backdoors must design them to avoid activation during training and evaluation. Since dat…

BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models

2024-06-24 · Yi Zeng, Weiyu Sun, Tran Ngoc Huynh, Dawn Song 외

Safety backdoor attacks in large language models (LLMs) enable the stealthy triggering of unsafe behaviors while evading detection during normal interactions. The high dimensionality of potential triggers in the token sp…

Code Generation

Propaganda via AI? A Study on Semantic Backdoors in Large Language Models

2025-04-15 · Nay Myat Min, Long H. Pham, Yige Li, Jun Sun

Large language models (LLMs) demonstrate remarkable performance across myriad language tasks, yet they remain vulnerable to backdoor attacks, where adversaries implant hidden triggers that systematically manipulate model…

Detecting and Eliminating Neural Network Backdoors Through Active Paths with Application to Intrusion Detection

2026-03-11 · Eirik Høyheim, Magnus Wiik Eckhoff, Gudmund Grov, Robert Flood 외 arxiv

Machine learning backdoors have the property that the machine learning model should work as expected on normal inputs, but when the input contains a specific $\textit{trigger}$, it behaves as the attacker desires. Detect…

Intrusion Detection

Beyond Corner Patches: Semantics-Aware Backdoor Attack in Federated Learning

2026-03-31 · Kavindu Herath, Joshua Zhao, Saurabh Bagchi arxiv

Backdoor attacks on federated learning (FL) are most often evaluated with synthetic corner patches or out-of-distribution (OOD) patterns that are unlikely to arise in practice. In this paper, we revisit the backdoor thre…

Traffic Sign RecognitionFederated Learning