paper-with-me

홈 › Papers

Learning to Give Checkable Answers with Prover-Verifier Games

2021-08-27 · Cem Anil, Guodong Zhang, Yuhuai Wu, Roger Grosse

Our ability to know when to trust the decisions made by machine learning systems has not kept up with the staggering improvements in their performance, limiting their applicability in high-stakes domains. We introduce Prover-Verifier Games (PVGs), a game-theoretic framework to encourage learning agents to solve decision problems in a verifiable manner. The PVG consists of two learners with competing objectives: a trusted verifier network tries to choose the correct answer, and a more powerful but untrusted prover network attempts to persuade the verifier of a particular answer, regardless of its correctness. The goal is for a reliable justification protocol to emerge from this game. We analyze variants of the framework, including simultaneous and sequential games, and narrow the space down to a subset of games which provably have the desired equilibria. We develop instantiations of the PVG for two algorithmic tasks, and show that in practice, the verifier learns a robust decision rule that is able to receive useful and reliable information from an untrusted prover. Importantly, the protocol still works even when the verifier is frozen and the prover's messages are directly optimized to convince the verifier.

📄 PDF Abstract BibTeX arXiv:2108.12099

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mitigating Legibility Tax with Decoupled Prover-Verifier Games

2026-02-26 · Yegon Kim, Juho Lee arxiv

As large language models become increasingly capable, it is critical that their outputs can be easily checked by less capable systems. Prover-verifier games can be used to improve checkability of model outputs, but displ…

Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction

2026-05-24 · João Sedoc, Baotong Zhang, Dean Foster arxiv

Reliably knowing when a language model is correct is almost as important as being correct. We introduce prover-verifier deliberation (PVD), an inference-time protocol grounded in interactive proof theory, as a mechanism …

Prover-Verifier Games improve legibility of LLM outputs

2024-07-18 · Jan Hendrik Kirchner, Yining Chen, Harri Edwards, Jan Leike 외

One way to increase confidence in the outputs of Large Language Models (LLMs) is to support them with reasoning that is clear and easy to check -- a property we call legibility. We study legibility in the context of solv…

Math

Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings

2025-07-10 · Berkant Turan, Suhrab Asadulla, David Steinmann, Kristian Kersting 외 arxiv

While Prover-Verifier Games (PVGs) offer a promising path toward verifiability in nonlinear classification models, they have not yet been applied to complex inputs such as high-dimensional images. Conversely, expressive …

Neural Interactive Proofs

2024-12-12 · Lewis Hammond, Sam Adam-Day

We consider the problem of how a trusted, but computationally bounded agent (a 'verifier') can learn to interact with one or more powerful but untrusted agents ('provers') in order to solve a given task. More specificall…