paper-with-me

홈 › Papers

The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring

2026-04-17 · Jon-Paul Cacioli arxiv

We introduce a cross-domain behavioural assay of monitoring-control coupling in LLMs, grounded in the Nelson and Narens (1990) metacognitive framework and applying human psychometric methodology to LLM evaluation. The battery comprises 524 items across six cognitive domains (learning, metacognitive calibration, social cognition, attention, executive function, prospective regulation), each grounded in an established experimental paradigm. Tasks T1-T5 were pre-registered on OSF prior to data collection; T6 was added as an exploratory extension. After every forced-choice response, dual probes adapted from Koriat and Goldsmith (1996) ask the model to KEEP or WITHDRAW its answer and to BET or decline. The critical metric is the withdraw delta: the difference in withdrawal rate between incorrect and correct items. Applied to 20 frontier LLMs (10,480 evaluations), the battery discriminates three profiles consistent with the Nelson-Narens architecture: blanket confidence, blanket withdrawal, and selective sensitivity. Accuracy rank and metacognitive sensitivity rank are largely inverted. Retrospective monitoring and prospective regulation appear dissociable (r = .17, 95% CI wide given n=20; exemplar-based evidence is the primary support). Scaling on metacognitive calibration is architecture-dependent: monotonically decreasing (Qwen), monotonically increasing (GPT-5.4), or flat (Gemma). Behavioural findings converge structurally with an independent Type-2 SDT approach, providing preliminary cross-method construct validity. All items, data, and code: https://github.com/synthiumjp/metacognitive-monitoring-battery.

📄 PDF Abstract BibTeX arXiv:2604.15702

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas

2026-04-21 · Jon-Paul Cacioli arxiv

Aggregate metacognitive quality scores mask within-model variation across MMLU benchmark domains. We administered 1,500 MMLU items (250 per domain, under an a priori six-domain grouping) to 33 frontier LLMs from eight mo…

What Is Going through Your Mind? Metacognitive Events Classification in Human-Agent Interactions

2022-06-01 · ISA (LREC) 2022 6 · Hafiza Erum Manzoor, Volha Petukhova

For an agent, either human or artificial, to show intelligent interactive behaviour implies assessments of the reliability of own and others’ thoughts, feelings and beliefs. Agents capable of these robust evaluations are…

Decision Makingvalid

LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling

2026-05-13 · Qi Cao, Yufan Wang, Peijia Qin, Shuhao Zhang 외 arxiv

Large language models (LLMs) often expose useful signals of self-monitoring: before solving a problem, they can estimate whether they are likely to succeed, and after solving it, they can judge whether their answer is li…

Multimodal Reasoning

Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning

2026-04-19 · Ziqing Zhuang, Linhai Zhang, Jiasheng Si, Deyu Zhou 외 arxiv

Large language models (LLMs) have demonstrated strong reasoning capabilities, and as existing approaches for enhancing LLM reasoning continue to mature, increasing attention has shifted toward meta-reasoning as a promisi…

Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models

2026-03-16 · Lit Sin Tan, Junzhe Chen, Xiaolong Fu, Lichen Ma 외 arxiv

Existing test-time scaling (TTS) methods for unified multimodal models (UMMs) in text-to-image (T2I) generation primarily rely on search or sampling strategies that produce only instance-level improvements, limiting the …

Reinforcement Learning