paper-with-me

홈 › Papers

Multidimensional Rubric-oriented Reward Model Learning via Geometric Projection Reference Constraints

2025-11-20 · Yongnan Jin, Xurui Li, Feng Cao, Liucun Gao, Juanjuan Yao arxiv

The integration of large language models (LLMs) into medical practice offers transformative potential, yet their real-world clinical applicability remains constrained by critical alignment issues: (1) a misalignment between static evaluation benchmarks and the dynamic cognitive demands of clinical practice, (2) challenges in adapting to continuously evolving, multi-source medical standards, and (3) the limited capacity of conventional reward models to reflect nuanced, multi-dimensional medical quality criteria. To overcome these limitations, we introduce MR-RML (Multidimensional Rubric-oriented Reward Model Learning) with GPRC (Geometric Projection Reference Constraints)-a novel alignment framework that structured medical standards into a multi-perspective matrix to guide both data generation and model optimization. Our approach introduces three key innovations: (1) a medical standard system that embeds domain-specific guidelines throughout the training pipeline; (2) an independent multi-dimensional reward model that decomposes evaluation criteria, transitioning from rule-based or LLM-based scoring to internalized reward modeling for better evaluation performance; and (3) geometric projection reference constraints that translate clinical cognitive logic into mathematical regularization, aligning scoring gradients with clinical reasoning and facilitating training with synthetically generated data. Extensive evaluations on the authoritative medical benchmark Healthbench demonstrate that our method significantly boosts the performance of the base Qwen-32B model, with improvements of 45% on the full subset and 85% on the hard subset. It achieves state-of-the-art results among open-source LLMs, scoring 62.7 (full) and 44.7 (hard), while also surpassing the majority of closed-source models.

📄 PDF Abstract BibTeX arXiv:2511.16139

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards

2026-05-26 · Yu Huang, Zihua Zhao, Zhaoxin Huan, Wanli Gu 외 arxiv

The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical dilemma inherent in this training paradig…

Reinforcement Learning

Thinking with Spatial Code for Physical-World Video Reasoning

2026-03-05 · Jieneng Chen, Wenxin Ma, Ruisheng Yuan, Yunzhi Zhang 외 arxiv

We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical-world visual question answering. We highlight the empirical finding that o…

Visual Question AnsweringReinforcement Learning

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

2026-08-26 · Shulin Tian, Minglun Li, Yuhao Dong, Hao Ding 외 arxiv

Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, or intermediate inference can undermine an otherwise plausible response…

Reinforcement LearningInstruction Following

Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards

2025-10-09 · Youliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan 외 arxiv

In this paper, we observe that current models are susceptible to reward hacking, leading to a substantial overestimation of a model's reasoning ability. This is evidenced by a high incidence of false positives-solutions …

Mathematical Reasoning

Multiclass threshold-based classification

2025-05-16 · Francesco Marchetti, Edoardo Legnaro, Sabrina Guastavino

In this paper, we introduce a threshold-based framework for multiclass classification that generalizes the standard argmax rule. This is done by replacing the probabilistic interpretation of softmax outputs with a geomet…

Classification