paper-with-me

홈 › Papers

RSD: Moving Local Triangular Charts for Auditing Language-Model Hidden States

2026-05-17 · Seungmin Jin arxiv

We study Relational Semantic Decomposition, abbreviated as RSD, as a moving local triangular chart audit for language-model hidden states. For repeated occurrences of one target word, RSD fits a shared three-anchor membership chart $S_t$ at layer or token-time $t$. The hidden-state channel uses $X_t\approx S_tC_t$; the invariant readout $M_t=S_tS_t^\top$ is the induced occurrence co-membership relation, and $R_t=X_t-S_tC_t$ records what the fitted root chart leaves outside the chart. The broader joint audit reuses the same membership chart for relation data, $A_t\approx S_tB_tS_t^\top$, such as an attention-derived occurrence relation. The current GPT-2 evidence is the $X$-channel hidden-state audit with Word-in-Context labels used as an external same-sense versus different-sense reference relation. On full WiC train, the root chart passes 16 of 53 eligible target words; this is audit coverage, not GPT-2 task accuracy. Token-time and pair-level diagnostics show the main regimes: \texttt{make} and \texttt{break} align at the target state, \texttt{drive} and \texttt{stay} improve after right context in small-count exploratory cases, and \texttt{play} remains a localized root-chart failure whose final same-sense pairs are not closer and have larger residual discrepancy. The resulting claim is diagnostic: RSD reports where a sense relation is visible in root co-membership and which failures become residual branch candidates or attention-channel obligations.

📄 PDF Abstract BibTeX arXiv:2605.17482

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Approximation of Functions over Manifolds: A Moving Least-Squares Approach

2017-11-02 · Barak Sober, Yariv Aizenbud, David Levin

We present an algorithm for approximating a function defined over a $d$-dimensional manifold utilizing only noisy function values at locations sampled from the manifold with noise. To produce the approximation we do not …

Dimensionality Reduction

RECAST: Interactive Auditing of Automatic Toxicity Detection Models

2020-01-07 · Austin P. Wright, Omar Shaikh, Haekyu Park, Will Epperson 외

As toxic language becomes nearly pervasive online, there has been increasing interest in leveraging the advancements in natural language processing (NLP), from very large transformer models to automatically detecting and…

Adversarial RobustnessFairness

EvidenceLens: A Claim-Evidence Matrix for Auditing Financial Question Answering

2026-06-19 · Fengchen Gu, Xiaotian Ren, Zhengyong Jiang, Zhilu Zhang 외 arxiv

Large language models are increasingly used to answer questions over annual reports, earnings decks, and analyst notes, yet their outputs remain difficult to verify in high-stakes financial workflows. A fluent answer can…

Question Answering

Martingale Doppelgänger-Eval: An Identification Framework for Auditing Candlestick Understanding in Vision-Language Models

2026-06-16 · Ziyao Wang arxiv

We introduce Martingale Doppelgänger-Eval, a public shadow-market benchmark for auditing whether vision-language models (VLMs) use candlestick evidence rather than extrapolate past trends. The central difficulty is ident…

Classifications of Single-input Lower Triangular Forms

2022-08-15 · Duan Zhang, Ying Sun

The purposes of this paper are to classify lower triangular forms and to determine under what conditions a nonlinear system is equivalent to a specific type of lower triangular forms. According to the least multi-indices…

FormVocal Bursts Type Prediction