paper-with-me

홈 › Papers

Thermodynamic Signatures of Reasoning: Free-Energy and Spectral-Form-Factor Diagnostics for Hallucination Detection in Large Language Models

2026-06-17 · Salim Khazem arxiv

Hallucination detection in large language models (LLMs) is deployment-critical, and recent work shows that the spectrum of attention-derived graph Laplacians carries strong signal about reasoning quality. Prior spectral diagnostics, however, summarize the Laplacian spectrum by a handful of eigenvalues or hand-picked scalars, leaving most of its structure unused. We propose Free-Energy Signatures (Fes), a spectral descriptor that treats each layer's attention Laplacian as a Hamiltonian and extracts its thermodynamic potentials partition function, free energy, spectral entropy, heat capacity together with the random-matrix-theory (RMT) spectral form factor. We prove three results: (i)~Lipschitz stability of Fes under attention perturbation; (ii)~an expressiveness result showing that Fes enriches finite spectral summaries and approximates moment-derived spectral functionals under explicit regularity and grid-resolution assumptions; and (iii)~a finite-sample PAC bound on the AUROC of a training-free detector built from Fes. Empirically, across six open-weight LLMs and six benchmarks, a lightweight probe on Fes descriptors achieves the strongest aggregate AUROC among attention-spectral baselines, improving over LapEig by $+6.5$ AUROC points and over GoR-4 by $+2.4$ points on average, while requiring no update to the underlying LLM. In the fully unsupervised setting, an RMT-deviation score achieves mean AUROC $0.71$, providing a label-free but weaker detector. A complementary RMT analysis shows that correct generations exhibit more Wigner-Dyson like spectral statistics, whereas hallucinations exhibit more Poisson-like statistics. The anonymized code and config are provided in the supplementary material.

📄 PDF Abstract BibTeX arXiv:2606.19404

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Graph Signal Processing Framework for Hallucination Detection in Large Language Models

2025-10-21 · Valentin Noël arxiv

Large language models achieve impressive results but distinguishing factual reasoning from hallucinations remains challenging. We propose a spectral analysis framework that models transformer layers as dynamic graphs ind…

Universal Thermodynamic Interatomic Potentials for Crystalline Materials

2026-08-14 · Juno Nam, Bowen Deng, Xiaochen Du, Luis Barroso-Luque 외 arxiv

Free energies govern solid-state phase stability, yet computational materials discovery still relies largely on ground-state energies because free energy calculations require ensemble averages. We introduce the thermodyn…

Guessing the upper bound free-energy difference between native-like structures

2018-11-21

Use of a combination of statistical thermodynamics and the Gershgorin theorem enable us to guess, in the thermodynamic limit, a plausible value for the upper bound free-energy difference between native-like structures of…

Thermodynamics-Consistent Graph Neural Networks

2024-07-08 · Jan G. Rittig, Alexander Mitsos

We propose excess Gibbs free energy graph neural networks (GE-GNNs) for predicting composition-dependent activity coefficients of binary mixtures. The GE-GNN architecture ensures thermodynamic consistency by predicting t…

Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning

2026-01-02 · Valentin Noël arxiv

Verifying whether a language model is genuinely reasoning or pattern-matching remains an open problem: learned verifiers are expensive, and output-based heuristics are brittle. We show that valid mathematical reasoning i…

Mathematical Reasoning