paper-with-me

Papers

Counterfactual Evaluation of Peer-Review Assignment Policies

2023-05-27 · NeurIPS 2023 11 · Martin Saveski, Steven Jecmen, Nihar B. Shah, Johan Ugander

Peer review assignment algorithms aim to match research papers to suitable expert reviewers, working to maximize the quality of the resulting reviews. A key challenge in designing effective assignment policies is evaluating how changes to the assignment algorithm map to changes in review quality. In this work, we leverage recently proposed policies that introduce randomness in peer-review assignment--in order to mitigate fraud--as a valuable opportunity to evaluate counterfactual assignment policies. Specifically, we exploit how such randomized assignments provide a positive probability of observing the reviews of many assignment policies of interest. To address challenges in applying standard off-policy evaluation methods, such as violations of positivity, we introduce novel methods for partial identification based on monotonicity and Lipschitz smoothness assumptions for the mapping between reviewer-paper covariates and outcomes. We apply our methods to peer-review data from two computer science venues: the TPDP'21 workshop (95 papers and 35 reviewers) and the AAAI'22 conference (8,450 papers and 3,145 reviewers). We consider estimates of (i) the effect on review quality when changing weights in the assignment algorithm, e.g., weighting reviewers' bids vs. textual similarity (between the review's past papers and the submission), and (ii) the "cost of randomization", capturing the difference in expected quality between the perturbed and unperturbed optimal match. We find that placing higher weight on text similarity results in higher review quality and that introducing randomization in the reviewer-paper assignment only marginally reduces the review quality. Our methods for partial identification may be of independent interest, while our off-policy approach can likely find use evaluating a broad class of algorithmic matching systems.

📄 PDF Abstract BibTeX arXiv:2305.17339

Code (1)

msaveski/counterfactual-peer-review 공식 구현

Tasks

counterfactualOff-policy evaluationtext similarity

Similar Papers 제목 키워드 기반

CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review

2026-09-04 · Jicheng Zhou, Kemou Li, Kahim Wong, Zheyuan Li 외 arxiv

Recent reports during the AAAI-27 review cycle highlight the risk of reviewers coordinating bids for reciprocal assignment advantage. Prior work treats bidding, reviewer assignment, and review manipulation as separate st…

Enhancing Peer Review in Astronomy: A Machine Learning and Optimization Approach to Reviewer Assignments for ALMA

2024-10-13 · John M. Carpenter, Andrea Corvillón, Nihar B. Shah

The increasing volume of papers and proposals that undergo peer review emphasizes the pressing need for greater automation to effectively manage the growing scale. In this study, we present the deployment and evaluation …

Astronomy

Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework

2025-08-29 · Nils Dycke, Iryna Gurevych arxiv

Large Language Models (LLMs) have great potential to accelerate and support scholarly peer review and are increasingly used as fully automatic review generators (ARGs). However, potential biases and systematic errors may…

Strategyproofing Peer Assessment via Partitioning: The Price in Terms of Evaluators' Expertise

2022-01-25 · Komal Dhull, Steven Jecmen, Pravesh Kothari, Nihar B. Shah

Strategic behavior is a fundamental problem in a variety of real-world applications that require some form of peer assessment, such as peer grading of homeworks, grant proposal review, conference peer review of scientifi…

PeerReview4All: Fair and Accurate Reviewer Assignment in Peer Review

2018-06-16 · Ivan Stelmakh, Nihar B. Shah, Aarti Singh

We consider the problem of automated assignment of papers to reviewers in conference peer review, with a focus on fairness and statistical accuracy. Our fairness objective is to maximize the review quality of the most di…

Fairness