Judgment-Grounded Expansion for Peer Review Generation
Automatic review generation is a promising direction for accelerating scientific progress. While most work adopts an end-to-end setup, its fully automated nature makes it less suitable for settings that demand accountability. To better balance automation and accountability, we formalize judgment-grounded expansion, a human-AI collaboration mode where a reviewer provides an evaluative claim and the system expands it into review comment candidate(s). We model it as a structured generate-check-refine process and conduct a user study to collect human-model interaction data. We study two practical challenges for judgment-grounded expansion: scalable evaluation and candidate set curation. We develop methods to simulate the process for large-scale evaluation, and show that conformal prediction is well suited to balancing candidate set size and target coverage. Our work establishes judgment-grounded expansion as a concrete task and provides empirical and methodological foundations for the design of future collaborative review generation systems.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Position on LLM-Assisted Peer Review: Addressing Reviewer Gap through Mentoring and Feedback
The rapid expansion of AI research has intensified the Reviewer Gap, threatening the peer-review sustainability and perpetuating a cycle of low-quality evaluations. This position paper critiques existing LLM approaches t…
PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
Evaluating large language models typically relies on human-authored benchmarks, reference answers, and human or single-model judgments, approaches that scale poorly, become quickly outdated, and mismatch open-world deplo…
Deep Transfer Learning Based Peer Review Aggregation and Meta-review Generation for Scientific Articles
Peer review is the quality assessment of a manuscript by one or more peer experts. Papers are submitted by the authors to scientific venues, and these papers must be reviewed by peers or other authors. The meta-reviewers…
ArticlesReview GenerationTransfer LearningFirstPass: Grounding AI Scientific Judgment in Multi-Round Editorial Outcomes
AI systems for peer review fail on three fronts: they train on Computer Science and Machine Learning venues alone, ignore the iterative dialogue that validates science, and evaluate on stylistic mimicry rather than real …
Benchmarking LLMs' Judgments with No Gold Standard
We introduce the GEM (Generative Estimator for Mutual Information), an evaluation metric for assessing language generation by Large Language Models (LLMs), particularly in generating informative judgments, without the ne…
BenchmarkingMachine TranslationText Generation