paper-with-me

홈 › Papers

Towards a Shared Rubric for Dataset Annotation

2021-12-07 · Andrew Marc Greene

When arranging for third-party data annotation, it can be hard to compare how well the competing providers apply best practices to create high-quality datasets. This leads to a "race to the bottom," where competition based solely on price makes it hard for vendors to charge for high-quality annotation. We propose a voluntary rubric which can be used (a) as a scorecard to compare vendors' offerings, (b) to communicate our expectations of the vendors more clearly and consistently than today, (c) to justify the expense of choosing someone other than the lowest bidder, and (d) to encourage annotation providers to improve their practices.

📄 PDF Abstract BibTeX arXiv:2112.03867

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rubric Reliability and Annotation of Content and Argument in Source-Based Argument Essays

2019-08-01 · WS 2019 8 · Yanjun Gao, Alex Driban, Brennan Xavier McManus, Elena Musi 외

We present a unique dataset of student source-based argument essays to facilitate research on the relations between content, argumentation skills, and assessment. Two classroom writing assignments were given to college s…

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge

2026-05-28 · Zijie Wang, Eduardo Blanco arxiv

LLM-as-a-Judge is a scalable alternative to human evaluation, yet existing rubric-based methods rely on human-annotated data such as reference answers or expert-crafted rubrics. We propose to automatically generate fine-…

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

2026-05-11 · Gaotang Li, Bhavana Dalvi Mishra, Zifeng Wang, Jun Yan 외 arxiv

Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyond the regime of verifiable rewards. Their outputs lack ground-truth…

Reinforcement Learning

C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences

2026-04-15 · Akira Kawabata, Saku Sugawara arxiv

Rubric-augmented verification guides reward models with explicit evaluation criteria, yielding more reliable judgments than single-model verification. However, most existing methods require costly rubric annotations, lim…

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

2026-07-16 · Haocheng Yang, Licheng Pan, Xiaoxi Li, Zhichao Chen 외 arxiv

Rubrics provide structured, fine-grained signals for training and evaluating large language models (LLMs). Yet reliable query-specific rubrics are difficult to construct. Existing approaches often derive supervision from…