paper-with-me

홈 › Papers

Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction

2026-05-24 · João Sedoc, Baotong Zhang, Dean Foster arxiv

Reliably knowing when a language model is correct is almost as important as being correct. We introduce prover-verifier deliberation (PVD), an inference-time protocol grounded in interactive proof theory, as a mechanism for selective prediction: the protocol produces both an answer and a structured confidence verdict, allowing a system to report high-confidence answers while abstaining on uncertain cases. In each dialogue, a prover defends a candidate answer through checkable sub-claims while a verifier issues targeted challenges and returns \textsc{Accept}, \textsc{Challenge}, or \textsc{Reject}. Because frozen language models are imperfect provers and verifiers operating over a noisy channel, formal soundness and completeness guarantees do not transfer; instead, we characterize the protocol empirically through its coverage-precision behavior. Our main experiment uses Claude Sonnet 4.6 as prover and Claude Haiku 4.5 as verifier on GPQA Diamond. Questions accepted with no answer revision, which we call Accept + No Change (ANC), are reported as the high-confidence subset; we evaluate this subset by its precision and coverage. ANC separates reliable from unreliable answers, yielding a $\sim$30pp HC-Prec gap over the non-ANC complement. Robustness experiments with GPT and Gemini pairings show that high HC-Prec can transfer across model families, while verifier strictness and domain competence largely determine the size of the selection gap. On Humanity's Last Exam, weaker prover-verifier pairings can collapse or invert the ANC signal, illustrating a practical failure mode when the verifier operates outside its effective region. Comparisons with self-consistency, universal self-consistency, multi-agent debate, and Reflexion suggest that prover-verifier deliberation supplies a distinct argument-defensibility signal for selective prediction.

📄 PDF Abstract BibTeX arXiv:2605.25133

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning to Give Checkable Answers with Prover-Verifier Games

2021-08-27 · Cem Anil, Guodong Zhang, Yuhuai Wu, Roger Grosse

Our ability to know when to trust the decisions made by machine learning systems has not kept up with the staggering improvements in their performance, limiting their applicability in high-stakes domains. We introduce Pr…

Efficiently Verifiable Proofs of Data Attribution

2025-08-14 · Ari Karchmer, Martin Pawelczyk, Seth Neel arxiv

Data attribution methods aim to answer useful counterfactual questions like "what would a ML model's prediction be if it were trained on a different dataset?" However, estimation of data attribution models through techni…

Prover-Verifier Games improve legibility of LLM outputs

2024-07-18 · Jan Hendrik Kirchner, Yining Chen, Harri Edwards, Jan Leike 외

One way to increase confidence in the outputs of Large Language Models (LLMs) is to support them with reasoning that is clear and easy to check -- a property we call legibility. We study legibility in the context of solv…

Math

Interactive proofs for verifying (quantum) learning and testing

2024-10-31 · Matthias C. Caro, Jens Eisert, Marcel Hinsche, Marios Ioannou 외

We consider the problem of testing and learning from data in the presence of resource constraints, such as limited memory or weak data access, which place limitations on the efficiency and feasibility of testing or learn…

TensorCommitments: A Lightweight Verifiable Inference for Language Models

2026-02-13 · Oguzhan Baser, Elahe Sadeghi, Eric Wang, David Ribeiro Alves 외 arxiv

Most large language models (LLMs) run on external clouds: users send a prompt, pay for inference, and must trust that the remote GPU executes the LLM without any adversarial tampering. We critically ask how to achieve ve…