paper-with-me

홈 › Papers

Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay

2026-08-27 · Prateek Chhikara arxiv

Interactive language-model agents use confidence signals to decide whether to answer immediately, retrieve additional evidence (from memory or external knowledge), or defer. Yet confidence is usually evaluated in isolation, without measuring the trajectory-level consequences of the actions it triggers. We propose matched trajectory replay, a controlled protocol for comparing confidence-to-action mappings. The protocol holds candidate answer states, evidence points, budgets, and action costs fixed. We use it to compare raw verbalized confidence with post-hoc isotonic calibration in a multi-hop question-answering system using Mistral, GPT, and Qwen models on HotpotQA and MuSiQue datasets. At the same numerical commitment threshold, calibration changes which questions agents ultimately commit to answering. Across all six model-dataset pairs, it increases accuracy among committed answers by up to 41 percentage points. However, it can reduce coverage and increase retrieval use. Overall accuracy improves by up to 15 percentage points on HotpotQA but falls by up to 17 percentage points on MuSiQue. These effects reflect a shift to a more selective, lower-risk operating point, not improved answers or confidence ranking. A calibration map fitted before retrieval improves held-out calibration through retrieval depths one and two, but is worse than raw confidence at depth three for all three models. Additional evidence helps on average, but this aggregate effect does not establish whether confidence identifies which individual episodes will benefit from another retrieval. Taken together, these results show that calibration can make commitment risk interpretable, but it does not estimate the expected benefit of another retrieval. Retrieval therefore requires a separate value-of-information or utility estimate. Evaluations should report held-out calibration, risk-coverage, and retrieval cost.

📄 PDF Abstract BibTeX arXiv:2608.26846

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents

2026-06-19 · Chubin Zhang, Zhenglin Wan, Xingrui Yu, Pengfei Zhou 외 arxiv

Tool-augmented agents are typically evaluated by their gains under reliable external feedback. Yet these gains leave open a key counterfactual: when feedback is unreliable, would the agent be better off receiving no task…

Question AnsweringFact Verification

PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval

2025-12-21 · Pengxiang Ouyang, Qing Ma, Zheng Wang, Cong Bai arxiv

Remote sensing (RS) image-text retrieval faces significant challenges in real-world datasets due to the presence of Pseudo-Matched Pairs (PMPs), semantically mismatched or weakly aligned image-text pairs, which hinder th…

Text Retrieval

Utterance-level neural confidence measure for end-to-end children speech recognition

2021-09-16 · Wei Liu, Tan Lee

Confidence measure is a performance index of particular importance for automatic speech recognition (ASR) systems deployed in real-world scenarios. In the present study, utterance-level neural confidence measure (NCM) in…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Binary Classificationspeech-recognition+1

Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs

2026-07-28 · Chandan Kumar Sah, Xiaoli Lian, Li Zhang arxiv

Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unnecessary computation. We investigate whether verbalized confidence fro…

Question Answering

MIND: Unified Inquiry and Diagnosis RL with Criteria Grounded Clinical Supports for Psychiatric Consultation

2026-03-04 · Guoyi Li, Shihao Xu, Jiatong Ma, Yunyun Han 외 arxiv

Psychiatric consultation requires agents to elicit discriminative evidence, map uncertain narratives to diagnostic criteria, and decide when evidence suffices. Existing dialogue and retrieval-augmented systems condition …

Reinforcement Learning