paper-with-me

홈 › Papers

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering

2026-07-28 · Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele Cinà, Luca Oneto, Iacopo Masi, Fabio Roli arxiv

Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text encoders, and exported computation graphs are distributed by third parties and reused across downstream services. This reuse model creates a security-critical trust boundary: VLM deployments inherit not only learned parameters but also executable behavior encoded in shared model artifacts. In this paper, we show that a malicious provider can exploit this trust boundary by embedding architectural backdoors into VLM supply chains through representation steering. Our attack introduces dormant steering logic into the model architecture through a trigger-gated additive modification of an intermediate representation, without poisoning training data, controlling downstream fine-tuning, or modifying prompts at deployment time. When the trigger is absent, the modification reduces to zero and the model follows its normal computation, preserving clean utility. When the trigger is present, a steering direction shifts the internal representation toward an attacker-defined objective. We evaluate the attack across multiple VLM families and downstream tasks, including visual question answering, text-to-image generation, retrieval, and semantic response biasing. The results show that the proposed architectural steering backdoor compromises integrity, safety enforcement, and ranking fairness while preserving normal behavior on clean inputs. We further show that shared VLM artifacts can carry dormant steering logic against downstream services, and we propose an auditing defense that inspects the executable logic distributed with model artifacts rather than only their learned weights.

📄 PDF Abstract BibTeX arXiv:2607.25479

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringText-to-Image Generation

Similar Papers 제목 키워드 기반

Inference-Time Backdoors via Chat Templates: From LLM Supply Chains to Agentic System Compromise

2026-02-04 · Ariel Fogel, Omer Hofman, Eilon Cohen, Roman Vainshtein arxiv

Open-weight language models are increasingly used in production settings, raising new security challenges. One prominent threat is backdoor attacks, in which adversaries embed hidden behaviors that activate under specifi…

Architectural Neural Backdoors from First Principles

2024-02-10 · Harry Langford, Ilia Shumailov, Yiren Zhao, Robert Mullins 외

While previous research backdoored neural networks by changing their parameters, recent work uncovered a more insidious threat: backdoors embedded within the definition of the network's architecture. This involves inject…

Architectural Backdoors in Neural Networks

2022-06-15 · CVPR 2023 1 · Mikel Bober-Irizar, Ilia Shumailov, Yiren Zhao, Robert Mullins 외

Machine learning is vulnerable to adversarial manipulation. Previous literature has demonstrated that at the training stage attackers can manipulate data and data sampling procedures to control model behaviour. A common …

Inductive Bias

Towards Automatic Discovery of Cybercrime Supply Chains

2018-12-02 · Rasika Bhalerao, Maxwell Aliapoulios, Ilia Shumailov, Sadia Afroz 외

Cybercrime forums enable modern criminal entrepreneurs to collaborate with other criminals into increasingly efficient and sophisticated criminal endeavors. Understanding the connections between different products and se…

Neurosymbolic Feature Extraction for Identifying Forced Labor in Supply Chains

2025-07-09 · Zili Wang, Frank Montabon, Kristin Yvonne Rozier arxiv

Supply chain networks are complex systems that are challenging to analyze; this problem is exacerbated when there are illicit activities involved in the supply chain, such as counterfeit parts, forced labor, or human tra…