paper-with-me

홈 › Papers

RubricBench: Aligning Model-Generated Rubrics with Human Standards

2026-03-02 · Qiyuan Zhang, Junyi Zhou, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma arxiv

As Large Language Model (LLM) alignment evolves from simple completions to complex, highly sophisticated generation, Reward Models are increasingly shifting toward rubric-guided evaluation to mitigate surface-level biases. However, the community lacks a unified benchmark to assess this evaluation paradigm, as existing benchmarks lack both the discriminative complexity and the ground-truth rubric annotations required for rigorous analysis. To bridge this gap, we introduce RubricBench, a curated benchmark with 1,147 pairwise comparisons specifically designed to assess the reliability of rubric-based evaluation. Our construction employs a multi-dimensional filtration pipeline to target hard samples featuring nuanced input complexity and misleading surface bias, augmenting each with expert-annotated, atomic rubrics derived strictly from instructions. Comprehensive experiments reveal a substantial capability gap between human-annotated and model-generated rubrics, indicating that even state-of-the-art models struggle to autonomously specify valid evaluation criteria, lagging considerably behind human-guided performance.

📄 PDF Abstract BibTeX arXiv:2603.01562

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics

2026-06-06 · Mengyuan Sun, Yu Li, Zhuohao Yu, Shikun Zhang 외 arxiv

Rubric-based evaluation is a promising paradigm for judging large language model (LLM) outputs, yet self-generated rubrics lag human-annotated criteria on hard instances. We argue this discriminative gap reflects an obje…

ClinAlign: Scaling Healthcare Alignment from Clinician Preference

2026-02-10 · Shiwei Lyu, Xidong Wang, Lei Liu, Hao Zhu 외 arxiv

Although large language models (LLMs) demonstrate expert-level medical knowledge, aligning their open-ended outputs with fine-grained clinician preferences remains challenging. Existing methods often rely on coarse objec…

Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability

2025-06-02 · Genta Indra Winata, David Anugraha, Emmy Liu, Alham Fikri Aji 외

High-quality datasets are fundamental to training and evaluating machine learning models, yet their creation-especially with accurate human annotations-remains a significant challenge. Many dataset paper submissions lack…

DescriptiveSynthetic Data Generation

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

2026-05-17 · Kuei-Chun Kao, Daixuan Huo, Yuanhao Ban, Cho-Jui Hsieh arxiv

Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images according to prompt alignment and perceptual quality. Existing reward m…

Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring

2024-07-04 · Xuansheng Wu, Padmaja Pravin Saraf, Gyeonggeon Lee, Ehsan Latif 외

Large language models (LLMs) have demonstrated strong potential in performing automatic scoring for constructed response assessments. While constructed responses graded by humans are usually based on given grading rubric…

Logical Reasoning