paper-with-me

홈 › Papers

A Unified Representation Underlying the Judgment of Large Language Models

2025-10-31 · Yi-Long Lu, Jiajun Song, Wei Wang arxiv

A central architectural question for both biological and artificial intelligence is whether judgment relies on specialized modules or a unified, domain-general resource. While the discovery of decodable neural representations for distinct concepts in Large Language Models (LLMs) has suggested a modular architecture, whether these representations are truly independent systems remains an open question. Here we provide evidence for a convergent architecture for evaluative judgment. Across a range of LLMs, we find that diverse evaluative judgments are computed along a dominant dimension, which we term the Valence-Assent Axis (VAA). This axis jointly encodes subjective valence ("what is good") and the model's assent to factual claims ("what is true"). Through direct interventions, we demonstrate this axis drives a critical mechanism, which is identified as the subordination of reasoning: the VAA functions as a control signal that steers the generative process to construct a rationale consistent with its evaluative state, even at the cost of factual accuracy. Our discovery offers a mechanistic account for response bias and hallucination, revealing how an architecture that promotes coherent judgment can systematically undermine faithful reasoning.

📄 PDF Abstract BibTeX arXiv:2510.27328

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

YiSi - a Unified Semantic MT Quality Evaluation and Estimation Metric for Languages with Different Levels of Available Resources

2019-08-01 · WS 2019 8 · Chi-kiu Lo

We present YiSi, a unified automatic semantic machine translation quality evaluation and estimation metric for languages with different levels of available resources. Underneath the interface with different language reso…

Machine TranslationSemantic SimilaritySemantic Textual SimilarityTranslation

Predicting human similarity judgments with distributional models: The value of word associations.

2016-12-01 · COLING 2016 12 · Simon De Deyne, Amy Perfors, Daniel J Navarro

Most distributional lexico-semantic models derive their representations based on external language resources such as text corpora. In this study, we propose that internal language models, that are more closely aligned to…

Language ModelingLanguage ModellingSemantic Textual Similarity

Extracting low-dimensional psychological representations from convolutional neural networks

2020-05-29 · Aditi Jha, Joshua Peterson, Thomas L. Griffiths

Deep neural networks are increasingly being used in cognitive modeling as a means of deriving representations for complex stimuli such as images. While the predictive power of these networks is high, it is often not clea…

How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective

2025-04-10 · Qi Liu, Jiaxin Mao, Ji-Rong Wen

Recent studies have shown that large language models (LLMs) can assess relevance and support information retrieval (IR) tasks such as document ranking and relevance judgment generation. However, the internal mechanisms b…

Document RankingInformation Retrieval

Using Natural Language Explanations to Rescale Human Judgments

2023-05-24 · Manya Wadhwa, Jifan Chen, Junyi Jessy Li, Greg Durrett

The rise of large language models (LLMs) has brought a critical need for high-quality human-labeled data, particularly for processes like human feedback and evaluation. A common practice is to label data via consensus an…

Question Answering