paper-with-me

홈 › Papers

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

2026-05-03 · Xiao Jia arxiv

Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. NeuroState-Bench is a human-calibrated benchmark that operationalizes commitment integrity through benchmark-defined side-query probes rather than inferred hidden activations. The released inventory contains 144 deterministic tasks and 306 benchmark-defined side-query probes spanning eight cognitively motivated failure families, paired clean and distractor variants, and three difficulty bands. The main 32-profile evaluation contains a fixed 16-profile local subset and a matched 16-profile hosted large-model subset evaluated through the same benchmark pipeline. Human calibration uses the final merged reporting scope: 104 sampled task units, 216 raw annotations, and 108 adjudicated task rows, with weighted kappa = 0.977 and ICC(2,1) = 0.977. Empirically, task success and commitment integrity diverge across this expanded grid: the success leader is not the integrity leader, 31 of 32 profiles change rank when integrity replaces task success, and integrity rankings are more stable under distractor perturbation. The primary confidence-free score HCCIS-CORE reaches 0.8469 AUC and 0.6992 PR-AUC for post-probe diagnostic discrimination of terminal task failure; the legacy full heuristic variant HCCIS-FULL reaches 0.7997 AUC and 0.6410 PR-AUC. Probe accuracy and state drift achieve slightly higher ROC-AUC, 0.8587, and better Brier/ECE, while HCCIS-CORE has substantially higher point-estimate PR-AUC and remains more closely tied to the benchmark's intended construct. The exploratory neural-augmented variant HCCIS+N is weaker overall, and a randomized subspace control approaches chance. NeuroState-Bench therefore contributes a calibrated evaluation axis for exposing commitment failures over a broader model grid than the original local-only subset.

📄 PDF Abstract BibTeX arXiv:2605.01847

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality

2026-05-03 · Wen Luo, Guangyue Peng, Liang Wang, Nan Yang 외 arxiv

Large Reasoning Models achieve strong performance on complex tasks but remain prone to hallucinations, particularly in long-form generation where errors compound across reasoning steps. Existing approaches to improving f…

SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation

2026-05-08 · Tianfei Ren, Zhipeng Yan, Yiming Zhao, Zhen Fang 외 arxiv

While text-to-image models have made strong progress in visual fidelity, faithfully realizing complex visual intents remains challenging because many requirements must be tracked across grounding, generation, and verific…

Image Generation

Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems

2026-04-19 · Tianyi Huang, Samuel Xu, Jason Tansong Dang, Samuel Yan 외 arxiv

Agentic systems often fail not by being entirely wrong, but by being too precise: a response may be generally useful while particular claims exceed what the evidence supports. We study this failure mode as overcommitment…

StakeBench: Evaluating Language Understanding Grounded in Market Commitment

2026-05-25 · Yunhua Pei, Jingyu Hu, Yiwei Shi, Hongnan Ma 외 arxiv

Existing financial NLP benchmarks often rely on labels supplied by outside observers, measuring how language is perceived rather than what speakers have committed to in the market. We introduce StakeBench, an evaluation …

Action Anticipation

Analytic Abduction: Causal Decomposition and Governed Commitment for Human--AI Coordination

2026-07-16 · Remo Pareschi arxiv

Abductive reasoning operates in two directions. The synthetic mode builds explanations from available hypotheses; the analytic mode, conversely, identifies the latent factors whose interaction accounts for a complex obse…