paper-with-me

홈 › Papers

The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

2026-08-25 · Augusto Camargo arxiv

Evaluations of generative language models frequently interpret observable behavioral traits, such as political stance, brand inclination, and normative framing, as manifestations of model weights, post-training alignment, or prompting. This interpretation risks conflating a foundation model with the multi-layered production system through which its outputs are ultimately served. Modern inference stacks support runtime interventions capable of modifying generation while model parameters remain frozen. We examine inference-time framing bias: systematic runtime steering of generated text toward institutional, ideological, or commercial frames without requiring changes to the underlying model parameters. We formalize the Inference Attribution Problem and establish an observational non-identifiability result showing that, under black-box observation alone, behaviorally equivalent deployed systems may arise from structurally distinct combinations of model parameters and inference policies. Consequently, observed behavioral bias does not uniquely identify the architectural layer responsible for it. We further characterize Probability Placement as a deployment pattern in which undisclosed commercial influence is embedded within an ostensibly organic assistant response through systematic probability-mass reallocation, distinguishing it from explicit token-auction mechanisms for generative advertising. Finally, we discuss implications for behavioral auditing, inference provenance, confidential computing, cryptographic attestation, the EU AI Act, the Digital Services Act, and advertising-disclosure principles. We argue that governance of generative systems must increasingly distinguish between auditing a model and auditing the deployed system that ultimately speaks.

📄 PDF Abstract BibTeX arXiv:2608.24662

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Finding needles in a haystack: A Black-Box Approach to Invisible Watermark Detection

2024-03-23 · Minzhou Pan, Zhenting Wang, Xin Dong, Vikash Sehwag 외

In this paper, we propose WaterMark Detection (WMD), the first invisible watermark detection method under a black-box and annotation-free setting. WMD is capable of detecting arbitrary watermarks within a given reference…

Challenge or Empower: Revisiting Argumentation Quality in a News Editorial Corpus

2018-10-01 · CONLL 2018 10 · Roxanne El Baff, Henning Wachsmuth, Khalid Al-Khatib, Benno Stein

News editorials are said to shape public opinion, which makes them a powerful tool and an important source of political argumentation. However, rarely do editorials change anyone{'}s stance on an issue completely, nor do…

Argument Mining

Editorial Alignment: A Participatory Approach to Engaging Editorial Expertise in LLM-mediated Knowledge Dissemination

2026-06-18 · Simon Aagaard Enni, Malthe Stavning Erslev, Karl-Emil Kjær Bilstrup, Kristoffer Laigaard Nielbo arxiv

The emergence of LLM-driven information services is reshaping the conditions under which public knowledge institutions operate, threatening to absorb the editorial function these institutions exist to exercise. While LLM…

Asymptotic Security of Control Systems by Covert Reaction: Repeated Signaling Game with Undisclosed Belief

2020-03-25

This study investigates the relationship between resilience of control systems to attacks and the information available to malicious attackers. Specifically, it is shown that control systems are guaranteed to be secure i…

Persuasiveness of News Editorials depending on Ideology and Personality

2020-12-01 · COLING (PEOPLES) 2020 12 · Roxanne El Baff, Khalid Al Khatib, Benno Stein, Henning Wachsmuth

News editorials aim to shape the opinions of their readership and the general public on timely controversial issues. The impact of an editorial on the reader’s opinion does not only depend on its content and style, but a…

Persuasiveness