paper-with-me

홈 › Papers

Reward Modeling for Scientific Writing Evaluation

2026-01-16 · Furkan Şahinuç, Subhabrata Dutta, Iryna Gurevych arxiv

Scientific writing is an expert-domain task that demands deep domain knowledge, task-specific requirements and reasoning capabilities that leverage the domain knowledge to satisfy the task specifications. While scientific text generation has been widely studied, its evaluation remains a challenging and open problem. It is critical to develop models that can be reliably deployed for evaluating diverse open-ended scientific writing tasks while adhering to their distinct requirements. However, existing LLM-based judges and reward models are primarily optimized for general-purpose benchmarks with fixed scoring rubrics and evaluation criteria. Consequently, they often fail to reason over sparse knowledge of scientific domains when interpreting task-dependent and multi-faceted criteria. Moreover, fine-tuning for each individual task is costly and impractical for low-resource settings. To bridge these gaps, we propose cost-efficient, open-source reward models tailored for scientific writing evaluation. We introduce a two-stage training framework that initially optimizes scientific evaluation preferences and then refines reasoning capabilities. Our multi-aspect evaluation design and joint training across diverse tasks enable fine-grained assessment and robustness to dynamic criteria and scoring rubrics. Experimental analysis shows that our training regime strongly improves LLM-based scientific writing evaluation. Our models generalize effectively across tasks and to previously unseen scientific writing evaluation settings, allowing a single trained evaluator to be reused without task-specific retraining.

📄 PDF Abstract BibTeX arXiv:2601.11374

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks

2026-04-30 · Qingyu Ren, Tianjun Pan, Xingzhou Chen, Xuhong Wang arxiv

Large language models have achieved remarkable progress in text generation but still struggle with generative writing tasks. In terms of evaluation, existing benchmarks evaluate writing reward models coarsely and fail to…

Reinforcement LearningText Generation

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

2026-08-10 · Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng 외 hf

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserve…

AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation

2025-04-10 · Tuhin Chakrabarty, Philippe Laban, Chien-Sheng Wu

AI-generated text is proliferating across domains, from creative writing and journalism to marketing content and scientific articles. Models can follow user-provided instructions to generate coherent and grammatically co…

ArticlesMarketing

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

2025-08-26 · Jianxing Liao, Tian Zhang, Xiao Feng, Yusong Zhang 외 arxiv

Large language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emotional expression) and objective constrai…

Reinforcement LearningInstruction Following

Expert Preference-based Evaluation of Automated Related Work Generation

2025-08-11 · Furkan Şahinuç, Subhabrata Dutta, Iryna Gurevych arxiv

Expert domain writing, such as scientific writing, typically demands extensive domain knowledge. Although large language models (LLMs) show promising potential in this task, evaluating the quality of automatically genera…