paper-with-me

홈 › Papers

No-Knowledge Alarms for Misaligned LLMs-as-Judges

2025-09-10 · Andrés Corrada-Emmanuel arxiv

If we use LLMs as judges to evaluate the complex decisions of other LLMs, who or what monitors the judges? Infinite monitoring chains are inevitable whenever we do not know the ground truth of the decisions by experts and we do not want to trust them. One way to ameliorate our evaluation uncertainty is to exploit the use of logical consistency between disagreeing experts. By observing how LLM judges agree and disagree while grading other LLMs, we can compute the only possible evaluations of their grading ability. For example, if two LLM judges disagree on which tasks a third one completed correctly, they cannot both be 100\% correct in their judgments. This logic can be formalized as a Linear Programming problem in the space of integer response counts for any finite test. We use it here to develop no-knowledge alarms for misaligned LLM judges. The alarms can detect, with no false positives, that at least one member or more of an ensemble of judges are violating a user specified grading ability requirement.

📄 PDF Abstract BibTeX arXiv:2509.08593

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Set of Documents to Present to an Analyst?

2020-05-01 · LREC 2020 5 · Richard Schwartz, John Makhoul, Lee Tarlin, Damianos Karakos

We describe the human triage scenario envisioned in the Cross-Lingual Information Retrieval (CLIR) problem of the [REDUCT] Program. The overall goal is to maximize the quality of the set of documents that is given to a b…

Cross-Lingual Information RetrievalInformation RetrievalRetrieval

Logical Consistency Between Disagreeing Experts and Its Role in AI Safety

2025-10-01 · Andrés Corrada-Emmanuel arxiv

If two experts disagree on a test, we may conclude both cannot be 100 per cent correct. But if they completely agree, no possible evaluation can be excluded. This asymmetry in the utility of agreements versus disagreemen…

Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce

2026-08-14 · Zeyuan Li, Lukas Petersson, Alessandro Acquisti, Michiel A. Bakker arxiv

Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elic…

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

2024-12-07 · Haitao Li, Qian Dong, Junjie Chen, Huixue Su 외

The rapid advancement of Large Language Models (LLMs) has driven their expanding application across various fields. One of the most promising applications is their role as evaluators based on natural language responses, …

Judging LLMs on a Simplex

2025-05-28 · Patrick Vossler, Fan Xia, Yifan Mai, Jean Feng

Automated evaluation of free-form outputs from large language models (LLMs) is challenging because many distinct answers can be equally valid. A common practice is to use LLMs themselves as judges, but the theoretical pr…

Bayesian InferenceUncertainty Quantification