paper-with-me

홈 › Papers

Geometry-Aware Backdoor Attacks: Leveraging Curvature in Hyperbolic Embeddings

2025-10-07 · Ali Baheri arxiv

Non-Euclidean foundation models increasingly place representations in curved spaces such as hyperbolic geometry. We show that this geometry creates a boundary-driven asymmetry that backdoor triggers can exploit. Near the boundary, small input changes appear subtle to standard input-space detectors but produce disproportionately large shifts in the model's representation space. Our analysis formalizes this effect and also reveals a limitation for defenses: methods that act by pulling points inward along the radius can suppress such triggers, but only by sacrificing useful model sensitivity in that same direction. Building on these insights, we propose a simple geometry-adaptive trigger and evaluate it across tasks and architectures. Empirically, attack success increases toward the boundary, whereas conventional detectors weaken, mirroring the theoretical trends. Together, these results surface a geometry-specific vulnerability in non-Euclidean models and offer analysis-backed guidance for designing and understanding the limits of defenses.

📄 PDF Abstract BibTeX arXiv:2510.06397

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stealthy Patch-Wise Backdoor Attack in 3D Point Cloud via Curvature Awareness

2025-03-12 · Yu Feng, Dingxin Zhang, Runkai Zhao, Yong Xia 외

Backdoor attacks pose a severe threat to deep neural networks (DNN) by implanting hidden backdoors that can be activated with predefined triggers to manipulate model behaviors maliciously. Existing 3D point cloud backdoo…

Backdoor Attack

Efficient Backdoor Removal Through Natural Gradient Fine-tuning

2023-06-30 · Nazmul Karim, Abdullah Al Arafat, Umar Khalid, Zhishan Guo 외

The success of a deep neural network (DNN) heavily relies on the details of the training scheme; e.g., training data, architectures, hyper-parameters, etc. Recent backdoor attacks suggest that an adversary can take advan…

backdoor defense

Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models

2025-08-30 · Zuquan Peng, Jianming Fu, Lixin Zou, Li Zheng 외 arxiv

The use of unvetted third-party and internet data renders pre-trained models susceptible to backdoor attacks. Detecting backdoor samples is critical to prevent backdoor activation during inference or injection during tra…

Contributor-Aware Defenses Against Adversarial Backdoor Attacks

2022-05-28 · Glenn Dawson, Muhammad Umer, Robi Polikar

Deep neural networks for image classification are well-known to be vulnerable to adversarial attacks. One such attack that has garnered recent attention is the adversarial backdoor attack, which has demonstrated the capa…

Backdoor Attackimage-classificationImage Classification

Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models

2026-06-29 · Arash Raftari, Mehrdad Mahdavi, Nathan Blackthorn, Andrew Arash Mahyari arxiv

Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is present. In this work, we study post ho…