paper-with-me

홈 › Papers

Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Quality Evaluation

2026-06-19 · Weilu Xu, Yunzhi Shen, Xinye Wang, Ranfei Dang, Shujian Huang arxiv

Large language models (LLMs) have shown strong potential in fine-grained translation quality evaluation (QE), yet existing MQM-based approaches typically rely on fixed rubric configurations shared across all translation samples. However, translation instances often differ substantially in error complexity, ambiguity, and required evaluation granularity, making static rubric allocation suboptimal for span-level error detection. We find that larger MQM subtype spaces improve error coverage but also introduce more false positives, while different translation instances prefer different rubric granularities, suggesting that evaluation spaces should be allocated dynamically for each case. Motivated by these observations, we propose a case-specific dynamic rubric framework that adaptively constructs MQM evaluation spaces for individual translation instances. Unlike fully free-form rubric generation methods, our framework remains grounded in the predefined MQM taxonomy while dynamically selecting suitable subtype spaces and evaluation granularity for different cases. Experiments on WMT span-level QE benchmarks across multiple model scales demonstrate that the proposed framework consistently improves MCC and produces cleaner span-level error localization compared with static rubric settings. Our results suggest that combining structured MQM rubrics with case-specific adaptive allocation is an effective strategy for fine-grained LLM-based translation evaluation.

📄 PDF Abstract BibTeX arXiv:2606.21559

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

2026-07-14 · Beidi Luan, Rui Sun, Sinuo Wang, Yan Gu 외 arxiv

Deep research agents are increasingly used to produce long-form financial reports, yet large-scale evaluation remains bottlenecked by the need for human experts to define and execute high-quality rubrics. We address this…

Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters

2026-04-27 · Aaryan Shah, Andrew Hines, Alexia Downs, Denis Bajet 외 arxiv

Objective. Clinical AI documentation systems require evaluation methodologies that are clinically valid, economically viable, and sensitive to iterative changes. Methods requiring expert review per scoring instance are t…

The USMLE® Step 2 Clinical Skills Patient Note Corpus

2022-07-01 · NAACL 2022 7 · Victoria Yaneva, Janet Mee, Le Ha, Polina Harik 외

This paper presents a corpus of 43,985 clinical patient notes (PNs) written by 35,156 examinees during the high-stakes USMLE® Step 2 Clinical Skills examination. In this exam, examinees interact with standardized patient…

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

2026-07-16 · Haocheng Yang, Licheng Pan, Xiaoxi Li, Zhichao Chen 외 arxiv

Rubrics provide structured, fine-grained signals for training and evaluating large language models (LLMs). Yet reliable query-specific rubrics are difficult to construct. Existing approaches often derive supervision from…

Are LLMs Bad at Moral Reasoning?

2026-06-10 · Menghang Zhu, Seth Lazar arxiv

For highly capable AI systems to operate safely in dynamic, open-ended environments, they must be able to identify, understand, and respond to moral reasons for action, and constrain their behaviour accordingly. A growin…