paper-with-me

홈 › Papers

From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks

2026-04-30 · Qingyu Ren, Tianjun Pan, Xingzhou Chen, Xuhong Wang arxiv

Large language models have achieved remarkable progress in text generation but still struggle with generative writing tasks. In terms of evaluation, existing benchmarks evaluate writing reward models coarsely and fail to measure performance from the perspective of specific requirements. In terms of training, existing training methods either use LLM-as-a-judge approaches or train coarse-grained reward models, lacking fine-grained requirement-adherence reward modeling. To address these issues, we propose a fine-grained evaluation pipeline WEval for writing reward models and a fine-grained reinforcement learning training framework WRL. The evaluation data of WEval covers multiple task categories and requirement types, enabling systematic evaluation of writing reward models by measuring the correlation between the rankings of the reward model and gold rankings. WRL constructs positive and negative samples by selectively dropping instruction requirements, allowing for more precise reward model training. Experiments show that our models achieve substantial improvements across various writing benchmarks and exhibit strong generalization. The code and data are publicly available at \href{https://github.com/Rainier-rq1/From_Coarse_to_Fine}{https://github.com/Rainier-rq1/From\_Coarse\_to\_Fine}.

📄 PDF Abstract BibTeX arXiv:2604.27453

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningText Generation

Similar Papers 제목 키워드 기반

Writer-R1: Enhancing Generative Writing in LLMs via Memory-augmented Replay Policy Optimization

2026-03-16 · Jihao Zhao, Shuaishuai Zu, Zhiyuan Ji, Chunlai Zhou 외 arxiv

As a typical open-ended generation task, creative writing lacks verifiable reference answers, which has long constrained reward modeling and automatic evaluation due to high human annotation costs, evaluative bias, and c…

Reinforcement Learning

Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards

2025-05-30 · Ruipeng Jia, Yunyi Yang, Yongbo Gai, Kai Luo 외

Reinforcement learning with verifiable rewards (RLVR) has enabled large language models (LLMs) to achieve remarkable breakthroughs in reasoning tasks with objective ground-truth answers, such as mathematics and code gene…

Code Generation

Reward Modeling for Scientific Writing Evaluation

2026-01-16 · Furkan Şahinuç, Subhabrata Dutta, Iryna Gurevych arxiv

Scientific writing is an expert-domain task that demands deep domain knowledge, task-specific requirements and reasoning capabilities that leverage the domain knowledge to satisfy the task specifications. While scientifi…

Text Generation

Coarse-to-Fine Process Reward Modeling for Enhanced Mathematical Reasoning

2025-01-23 · Yulan Hu, Sheng Ouyang, Yong liu

Process reward model (PRM) is critical for mathematical reasoning tasks to assign rewards for each intermediate steps. The PRM requires constructing process-wise supervision data for training, which rely on chain-of-thou…

AttributeMathematical Reasoning

HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning

2025-11-19 · Qihao Yang, Xuelin Wang, Jiale Chen, Xuelian Dong 외 arxiv

Language acquisition is vital to revealing the nature of human language intelligence and has recently emerged as a promising perspective for improving the interpretability of large language models (LLMs). However, it is …

Language Acquisition