paper-with-me

Papers

LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components

2025-08-24 · Hikaru Tsujimura, Arush Tagade arxiv

Large Language Models (LLMs) often display overconfidence, presenting information with unwarranted certainty in high-stakes contexts. We investigate the internal basis of this behavior via mechanistic interpretability. Using open-sourced Llama 3.2 models fine-tuned on human annotated assertiveness datasets, we extract residual activations across all layers, and compute similarity metrics to localize assertive representations. Our analysis identifies layers most sensitive to assertiveness contrasts and reveals that high-assertive representations decompose into two orthogonal sub-components of emotional and logical clusters-paralleling the dual-route Elaboration Likelihood Model in Psychology. Steering vectors derived from these sub-components show distinct causal effects: emotional vectors broadly influence prediction accuracy, while logical vectors exert more localized effects. These findings provide mechanistic evidence for the multi-component structure of LLM assertiveness and highlight avenues for mitigating overconfident behavior.

📄 PDF Abstract BibTeX arXiv:2508.17182

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Corpus for Research on Deliberation and Debate

2012-05-01 · LREC 2012 5 · Marilyn Walker, Jean Fox Tree, Pranav Anand, Rob Abbott 외

Deliberative, argumentative discourse is an important component of opinion formation, belief revision, and knowledge discovery; it is a cornerstone of modern civil society. Argumentation is productively studied in branch…

Large Language Models Meet Biomedical Knowledge Graphs for Mechanistically Grounded Therapeutic Prioritization

2026-04-17 · Chih-Hsuan Wei, Chi-Ping Day, Zhizheng Wang, Christine C. Alewine 외 arxiv

Drug repurposing is often framed as a candidate identification task, but existing approaches provide limited guidance for distinguishing biologically plausible candidates from historically well-connected ones. Here we in…

Knowledge Graphs

I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue System

2026-06-10 · Zi Haur Pang, Yahui Fu, Koji Inoue, Tatsuya Kawahara arxiv

Emotional validation - explicitly acknowledging that a user's feelings make sense - has proven therapeutic value but has received little computational attention. Emotional validation in dialogue systems can be decomposed…

Response Generation

Human Feedback is not Gold Standard

2023-09-28 · Tom Hosking, Phil Blunsom, Max Bartolo

Human feedback has become the de facto standard for evaluating the performance of Large Language Models, and is increasingly being used as a training objective. However, it is not clear which properties of a generated ou…

Emotionally Charged, Logically Blurred: AI-driven Emotional Framing Impairs Human Fallacy Detection

2025-10-09 · Yanran Chen, Lynn Greschner, Roman Klinger, Michael Klenk 외 arxiv

Logical fallacies are common in public communication and can mislead audiences; fallacious arguments may still appear convincing despite lacking soundness, because convincingness is inherently subjective. We present the …

Logical Fallacies