paper-with-me

홈 › Papers

A Sharp Test for the Judge Leniency Design

2024-05-10 · Mohamed Coulibaly, Yu-Chin Hsu, Ismael Mourifié, Yuanyuan Wan

We propose a new specification test to assess the validity of the judge leniency design. We characterize a set of sharp testable implications, which exploit all the relevant information in the observed data distribution to detect violations of the judge leniency design assumptions. The proposed sharp test is asymptotically valid and consistent and will not make discordant recommendations. When the judge's leniency design assumptions are rejected, we propose a way to salvage the model using partial monotonicity and exclusion assumptions, under which a variant of the Local Instrumental Variable (LIV) estimand can recover the Marginal Treatment Effect. Simulation studies show our test outperforms existing non-sharp tests by significant margins. We apply our test to assess the validity of the judge leniency design using data from Stevenson (2018), and it rejects the validity for three crime categories: robbery, drug selling, and drug possession.

📄 PDF Abstract BibTeX arXiv:2405.06156

Code (0)

등록된 구현이 없습니다.

Tasks

valid

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

2026-06-30 · Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmad ElShiekh, Clayton W. Taylor 외 arxiv

The cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, high-stakes domains. This is particularly acute for Arabic sociolinguistic knowledge: credible grading requires n…

Evaluative Fingerprints: Stable and Systematic Differences in LLM Evaluator Behavior

2026-01-08 · Wajid Nasser arxiv

LLM-as-judge systems promise scalable, consistent evaluation. We find the opposite: judges are consistent, but not with each other; they are consistent with themselves. Across 3,240 evaluations (9 judges x 120 unique vid…

Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

2024-06-18 · Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan 외

Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating large language models (LLMs). However, th…

TriviaQA

Context Over Content: Exposing Evaluation Faking in Automated Judges

2026-04-16 · Manan Gupta, Inderjeet Nair, Lu Wang, Dhruv Kumar arxiv

The $\textit{LLM-as-a-judge}$ paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: that judges evaluate text strictly on its semantic content, impervio…

Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG

2026-07-12 · Sriram Selvam, Anneswa Ghosh arxiv

LLM-as-a-judge evaluation is widely used for retrieval-augmented generation (RAG), but reusing the same model family as both generator and judge makes self-leniency difficult to identify. We introduce Eval-Pair Matrix, a…