paper-with-me

홈 › Papers

Investigating Linguistic Steering: An Analysis of Adjectival Effects Across Large Language Model Architectures

2026-04-28 · Lars Malmqvist arxiv

Achieving reliable control of Large Language Models (LLMs) requires a precise, scalable understanding of how they interpret linguistic cues. We introduce a rigorous framework using Shapley values to quantify the steering effect of individual adjectives on model performance, moving beyond anecdotal heuristics to principled attribution. Applying this method to 100 adjectives across a diverse suite of models (including o3, gpt-4o-mini, phi-3, llama-3-70b, and deepseek-r1) on the MMLU benchmark, we uncover several critical findings for AI alignment. First, we find that a small subset of adjectives act as disproportionately powerful "levers," yet their effects are not universal. Cross-model analysis reveals a "family effect": models of a shared lineage exhibit correlated sensitivity profiles, while architecturally distinct models react in a largely uncorrelated manner, challenging the notion of a one-size-fits-all prompting strategy. Second, focused follow-up studies demonstrate that the steering direction of these powerful adjectives is not intrinsic but is highly contingent on their syntactic role and position within the prompt. For larger models like gpt-4o-mini, we provide the first quantitative evidence of strong, non-additive interaction effects where adjectives can synergistically amplify, antagonistically dampen, or even reverse each other's impact. In contrast, smaller models like phi-3 exhibit a more literal and less compositional response. These results suggest that as models scale, their interpretation of prompts becomes more sophisticated but also less predictable, posing a significant challenge for robustly steering model behavior and highlighting the need for compositional and model-specific alignment techniques.

📄 PDF Abstract BibTeX arXiv:2606.20572

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Clique-based Graphical Approach to Detect Interpretable Adjectival Senses in Hungarian

2022-10-01 · COLING (TextGraphs) 2022 10 · Enikő Héja, Noémi Ligeti-Nagy

The present paper introduces an ongoing research which aims to detect interpretable adjectival senses from monolingual corpora applying an unsupervised WSI approach. According to our expectations the findings of our inve…

Steering Prepositional Phrases in Language Models: A Case of with-headed Adjectival and Adverbial Complements in Gemma-2

2025-09-27 · Stefan Arnold, René Gröbner arxiv

Language Models, when generating prepositional phrases, must often decide for whether their complements functions as an instrumental adjunct (describing the verb adverbially) or an attributive modifier (enriching the nou…

Analysis of Similes in Serbian Literary Texts (1860-1920) using computational methods

2020-09-01 · CLIB 2020 9 · Cvetana Krstev, Jelena Jaćimović, Duško Vitas

Similes are rhetorical figures which play an important role in literary texts. This paper presents a finite-state methodology developed for the description of adjectival similes, which enables their retrieval and annotat…

RetrievalSpecificity

How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects

2025-10-08 · Leonardo Bertolazzi, Sandro Pezzelle, Raffaella Bernardi arxiv

Both humans and large language models (LLMs) exhibit content effects: biases in which the plausibility of the semantic content of a reasoning problem influences judgments regarding its logical validity. While this phenom…

PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting

2025-09-18 · Caitlin Cisar, Emily Sheffield, Joshua Drake, Alden Harrell 외 arxiv

Generative AI applications commonly leverage user personas as a steering mechanism for synthetic data generation, but reliance on natural language representations forces models to make unintended inferences about which a…

Synthetic Data Generation