paper-with-me

홈 › Papers

Training AI Co-Scientists Using Rubric Rewards

2025-12-29 · Shashwat Goel, Rishi Hazra, Dulhan Jayalath, Timon Willi, Parag Jain, William F. Shen, Ilias Leontiadis, Francesco Barbieri, Yoram Bachrach, Jonas Geiping, Chenxi Whitehouse arxiv

AI co-scientists are emerging as a tool to assist human researchers in achieving their research goals. A crucial feature of these AI co-scientists is the ability to generate a research plan given a set of aims and constraints. The plan may be used by researchers for brainstorming, or may even be implemented after further refinement. However, language models currently struggle to generate research plans that follow all constraints and implicit requirements. In this work, we study how to leverage the vast corpus of existing research papers to train language models that generate better research plans. We build a scalable, diverse training corpus by automatically extracting research goals and goal-specific grading rubrics from papers across several domains. We then train models for research plan generation via reinforcement learning with self-grading. A frozen copy of the initial policy acts as the grader during training, with the rubrics creating a generator-verifier gap that enables improvements without external human supervision. To validate this approach, we conduct a study with human experts for machine learning research goals, spanning 225 hours. The experts prefer plans generated by our finetuned Qwen3-30B-A3B model over the initial model for 70% of research goals, and approve 84% of the automatically extracted goal-specific grading rubrics. To assess generality, we also extend our approach to research goals from medical papers, and new arXiv preprints, evaluating with a jury of frontier models. Our finetuning yields 12-22% relative improvements and significant cross-domain generalization, proving effective even in problem settings like medical research where execution feedback is infeasible. Together, these findings demonstrate the potential of a scalable, automated training recipe as a step towards improving general AI co-scientists.

📄 PDF Abstract BibTeX arXiv:2512.23707

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningDomain Generalization

Similar Papers 제목 키워드 기반

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

2026-05-19 · Utkarsh Tyagi, Xingang Guo, MohammadHossein Rezaei, Daniel George 외 arxiv

Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many important model behaviors require satisfying several qualitative criteri…

Reinforcement Learning

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

2025-07-23 · Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding. However, extending it to real-world reasoning tasks is challe…

Reinforcement Learning

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

2026-08-03 · Fangxu Yu, Tao Feng, Dehai Min, Zinan Lin 외 hf

Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: …

Reinforcement Learning

Online Rubrics Elicitation from Pairwise Comparisons

2025-10-08 · MohammadHossein Rezaei, Robert Vacareanu, Zihao Wang, Clinton Wang 외 arxiv

Rubrics provide a flexible way to train LLMs on open-ended long-form answers where verifiable rewards are not applicable and human preferences provide coarse signals. Prior work shows that reinforcement learning with rub…

Reinforcement Learning

SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing

2026-02-24 · Yifei Xu, Guilherme Potje, Shivam Shandilya, Tiancheng Yuan 외 arxiv

Designing aligned and robust rewards for open-ended generation remains a key barrier to RL post-training. Rubrics provide structured, interpretable supervision, but scaling rubric construction is difficult: expert rubric…