paper-with-me

Papers

A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation

2026-08-14 · Dipankar Sarkar arxiv

Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute. Our position is that any such judge must be evaluated on four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. We release Principle-Bench, 168 cryptoasset financial-promotion scenarios mapped to two UK FCA principles, with paraphrase, adversarial keyword-stuffing, and boundary perturbations authored under a pre-registered rubric; the first benchmark covering all four axes for principle-based regulation. We also introduce Ceca (Calibrated Exemplar-Cluster Assessment): a calibrated, auditable assessor that emits exact per-exemplar counterfactual attributions. Across keyword counting, three sentence-transformer embedders, an open-weight LLM-judge, and a calibrated cascade, no method dominates all four axes. A 120B LLM-judge, strongest on benign inputs, loses 47 accuracy points (0.74 to 0.27) on keyword-stuffed Consumer Duty inputs: "compliance theatre." A second judge from a different model family agrees only at Cohen's kappa = 0.16 on that split, localising the failure to the model rather than the corpus. Any deployment-grade LLM-judge for principle-based regulation must report per-principle adversarial deception and post-hoc calibration alongside aggregate accuracy.

📄 PDF Abstract BibTeX arXiv:2608.14329

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Robustness

Similar Papers 제목 키워드 기반

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

2026-06-02 · Sourabrata Mukherjee, Hamna Hamna, Kalika Bali, Sunayana Sitaram arxiv

LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared signal or shared bias by measuring four geometric quantities on the …

Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

2026-05-27 · Richard J. Young, Gregory D. Moody arxiv

A general-purpose language model that answers a harmful question returns text; a coding model that complies with a malicious request can return a working weapon: a keylogger, ransomware, an exploit that runs as written. …

GSAR: Typed Grounding for Hallucination Detection and Recovery in Multi-Agent LLMs

2026-04-25 · Federico A. Kamelhar arxiv

Autonomous multi-agent LLM systems are increasingly deployed to investigate operational incidents and produce structured diagnostic reports. Their trustworthiness hinges on whether each claim is grounded in observed evid…

Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment

2026-02-27 · Dake Zhang, Mark D. Smucker, Charles L. A. Clarke arxiv

Many readers today struggle to assess the trustworthiness of online news because reliable reporting coexists with misinformation. The TREC 2025 DRAGUN (Detection, Retrieval, and Augmented Generation for Understanding New…

Question Generation

Social Meaning in Repeated Interactions

2020-06-01 · PaM 2020 6 · Elin McCready, Robert Henderson

Judgements about communicative agents evolve over the course of interactions both in how individuals are judged for testimonial reliability and for (ideological) trustworthiness. This paper combines a theory of social me…