paper-with-me

Papers

Cherry-pick Override: Unsafe Directional Commitment in LLM Judges under Mixed Evidence

2026-06-05 · Haoran Xu arxiv

LLM judges increasingly turn verdicts into system commitments. Under mixed evidence (claims with both supporting and refuting sources) this is unsafe: when the schema exposes CONFLICTING as the authorized non-directional verdict, returning SUPPORTS/REFUTES is an unauthorized directional commitment, a failure we name Cherry-pick Override (CCO). We define CCO under an explicit task contract and report it with a same-denominator diagnostic protocol paired with matched-coverage bootstrap and an apples-to-apples random-veto null. On AVeriTeC's Conflicting subset (N_C = 150), three-option judges return a directional verdict on more than 84% of mixed-evidence claims; under the typed schema, three-judge majority voting amplifies direction-on-conflict on AVeriTeC (0.887 vs. 0.840; 95% CI [+0.013, +0.080]) but does not replicate on VitaminC-Mixed. Walking an intervention ladder of common single-channel fixes (typed vocabulary, panel aggregation, confidence thresholding, validator-only filtering), each leaves a distinct residual failure: panel aggregation suppresses single-judge CONFLICTING dissent in 48% of CCO cases; the panel is well-calibrated for direction (ECE = 0.07 on pure-S/R) so confidence cannot operationally separate CCO from correct directional commits; validator-as-classifier nearly halves pure-evidence accuracy. A minimal two-channel reference probe reaches operating points neither single channel reaches; under the random-veto null its promotion to CONFLICTING is structurally targeted on AVeriTeC (empirical p < 1/2001) and weaker but in the same direction on VitaminC-Mixed, a selectivity result rather than a magnitude one. We argue for an external commitment-control layer that separates verdict generation from commitment authorization, using structural evidence and confidence as orthogonal channels and NO-COMMIT as a routed controller state.

📄 PDF Abstract BibTeX arXiv:2606.07834

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On Context-aware Detection of Cherry-picking in News Reporting

2024-01-11 · Israa Jaradat, Haiqi Zhang, Chengkai Li

Cherry-picking refers to the deliberate selection of evidence or facts that favor a particular viewpoint while ignoring or distorting evidence that supports an opposing perspective. Manually identifying cherry-picked sta…

On the Definition and Detection of Cherry-Picking in Counterfactual Explanations

2026-01-08 · James Hinns, Sofie Goethals, Stephan Van der Veeken, Theodoros Evgeniou 외 arxiv

Counterfactual explanations are widely used to communicate how inputs must change for a model to alter its prediction. For a single instance, many valid counterfactuals can exist, which leaves open the possibility for an…

On the existence of a cherry-picking sequence

2017-12-12

Recently, the minimum number of reticulation events that is required to simultaneously embed a collection P of rooted binary phylogenetic trees into a so-called temporal network has been characterized in terms of cherry-…

Cherry on the Cake: Fairness is NOT an Optimization Problem

2024-06-24 · Marco Favier, Toon Calders

In Fair AI literature, the practice of maliciously creating unfair models that nevertheless satisfy fairness constraints is known as "cherry-picking". A cherry-picking model is a model that makes mistakes on purpose, sel…

FairnessMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

CherryPicker: Semantic Skeletonization and Topological Reconstruction of Cherry Trees

2023-04-10 · Lukas Meyer, Andreas Gilson, Oliver Scholz, Marc Stamminger

In plant phenotyping, accurate trait extraction from 3D point clouds of trees is still an open problem. For automatic modeling and trait extraction of tree organs such as blossoms and fruits, the semantically segmented p…

Monocular ReconstructionPlant PhenotypingSemantic Segmentation