paper-with-me

홈 › Papers

Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report

2026-04-20 · Jon-Paul Cacioli arxiv

Clinical personality assessment screens response validity before interpreting substantive scales. LLM evaluation does not. We apply the validity scaling framework from the PAI and MMPI-3 to metacognitive probe data from 20 frontier models across 524 items. Six validity indices are operationalised: L (maintaining confidence on errors), K (betting on errors), F (withdrawing consensus-endorsed items), Fp (withdrawing correct answers), RBS (inverted monitoring), and TRIN (fixed responding). A tiered classification system identifies four models as construct-level invalid and two as elevated. Valid-profile models produce item-sensitive confidence (mean r = .18, 14 of 16 significant). Invalid-profile models do not (mean r = -.20, d = 2.17, p = .001). Chain-of-thought training produces two opposite response distortions. Two latent dimensions account for 94.6% of index variance. Companion papers extract a portable screening protocol (Cacioli, 2026e) and validate it against selective prediction (Cacioli, 2026f). All data and code: https://github.com/synthiumjp/validity-scaling-llm

📄 PDF Abstract BibTeX arXiv:2604.17707

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals

2026-04-20 · Jon-Paul Cacioli arxiv

LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence signal carries item-level information before building on it. We transf…

The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring

2026-04-17 · Jon-Paul Cacioli arxiv

We introduce a cross-domain behavioural assay of monitoring-control coupling in LLMs, grounded in the Nelson and Narens (1990) metacognitive framework and applying human psychometric methodology to LLM evaluation. The ba…

LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling

2026-05-13 · Qi Cao, Yufan Wang, Peijia Qin, Shuhao Zhang 외 arxiv

Large language models (LLMs) often expose useful signals of self-monitoring: before solving a problem, they can estimate whether they are likely to succeed, and after solving it, they can judge whether their answer is li…

Multimodal Reasoning

MetaCogAgent: A Metacognitive Multi-Agent LLM Framework with Self-Aware Task Delegation

2026-05-17 · Chenyu Wang, Yang Shu arxiv

Multi-agent large language model (LLM) systems have shown promise for solving complex tasks through agent collaboration. However, existing frameworks assign tasks based on predefined roles without considering whether an …

Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas

2026-04-21 · Jon-Paul Cacioli arxiv

Aggregate metacognitive quality scores mask within-model variation across MMLU benchmark domains. We administered 1,500 MMLU items (250 per domain, under an a priori six-domain grouping) to 33 frontier LLMs from eight mo…