paper-with-me

홈 › Papers

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

2026-07-09 · Raphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana, Christopher J. Earls, Srikar Varadaraj, Eric Ho arxiv

Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a more direct window into both. Working with Eternis-Forecaster 8B on OpenForesight, we train representation-pooling probes on intermediate activations and find they achieve substantially better calibration; a result that also holds for GLM-4.7-Flash and GLM-4.5-Air. We then assess CoT faithfulness through evidence ablation and diversionary injection: removing an influential source in the prompt often changes the model's forecast while leaving the reasoning trace untouched. The same probes function as lie detectors: their activations track behavioral shifts far better than the reasoning trace does, and they also predict the direction of change in 84% of cases, including when the CoT conceals the perturbation's influence. Finally, forced answering reveals that forecasts are largely fixed before reasoning begins: a single pre-reasoning pass recovers the committed answer and confidence, and routing questions by the spread of this pre-set answer distribution saves 30-47% of generated tokens, with no loss of accuracy. Together, these results establish probing internal representations as a practical tool for calibrating, auditing, and triaging language model forecasters and reasoning models more broadly.

📄 PDF Abstract BibTeX arXiv:2607.08046

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 111
arxivsub/arXivSub_daily_arxiv ★ 2

Similar Papers 제목 키워드 기반

Concept Probing: Where to Find Human-Defined Concepts (Extended Version)

2025-07-24 · Manuel de Sousa Ribeiro, Afonso Leote, João Leite arxiv

Concept probing has recently gained popularity as a way for humans to peek into what is encoded within artificial neural networks. In concept probing, additional classifiers are trained to map the internal representation…

What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators

2026-03-23 · Xinyu Zhang arxiv

World models learn to simulate environment dynamics from experience, enabling sample-efficient reinforcement learning. But what do these models actually represent internally? We apply interpretability techniques--includi…

Reinforcement Learning

Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models

2026-06-07 · Lucrezia Laraspata, Giovanna Castellano, Gennaro Vessio arxiv

Hallucinations -- factually incorrect or unverifiable outputs -- remain one of the most challenging limitations of Large Language Models (LLMs), especially in knowledge-intensive tasks. One proposed explanation is intern…

Dense Passage Retrieval: Is it Retrieving?

2024-02-16 · Benjamin Reichman, Larry Heck

Dense passage retrieval (DPR) is the first step in the retrieval augmented generation (RAG) paradigm for improving the performance of large language models (LLM). DPR fine-tunes pre-trained networks to enhance the alignm…

Model EditingPassage RetrievalRAGRetrieval+1

Probing Neural Combinatorial Optimization Models

2025-10-25 · Zhiqin Zhang, Yining Ma, Zhiguang Cao, Hoong Chuin Lau arxiv

Neural combinatorial optimization (NCO) has achieved remarkable performance, yet its learned model representations and decision rationale remain a black box. This impedes both academic research and practical deployment, …