paper-with-me

Papers

DocReward: A Document Reward Model for Structuring and Stylizing

2025-10-13 · Junpeng Liu, Yuzhong Zhao, Bowen Cao, Jiayu Ding, Yilin Jia, Tengchao Lv, Yupan Huang, Wenshan Wu, Shaohan Huang, Nan Yang, Li Dong, Lei Cui, Tao Ge, Xun Wang, Huitian Jiao, Sun Mao, FNU Kartik, Si-Qing Chen, Wai Lam, Furu Wei arxiv

Recent agentic workflows automate professional document generation but focus narrowly on textual quality, overlooking structural and stylistic professionalism, which is equally critical for readability. This gap stems mainly from a lack of effective reward models capable of guiding agents toward producing documents with high structural and stylistic professionalism. We introduce DocReward, a document reward model that evaluates documents based on their structure and style. To achieve this, we propose a textual-quality-agnostic framework that ensures assessments are not confounded by content quality, and construct DocPair, a dataset of 117K paired documents covering 32 domains and 267 types. Each pair shares identical content but differs in structural and stylistic professionalism. DocReward is trained using the Bradley-Terry loss. On a manually annotated benchmark, DocReward outperforms GPT-5 by 14.6 percentage points in the same setting. Reinforcement learning experiments further show that DocReward effectively guides agents toward generating documents with consistently higher structural and stylistic professionalism, highlighting its practical utility.

📄 PDF Abstract BibTeX arXiv:2510.11391

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Seg2Act: Global Context-aware Action Generation for Document Logical Structuring

2024-10-09 · Zichao Li, Shaojie He, Meng Liao, Xuanang Chen 외

Document logical structuring aims to extract the underlying hierarchical structure of documents, which is crucial for document intelligence. Traditional approaches often fall short in handling the complexity and the vari…

Action GenerationTransfer Learning

Structuring an unordered text document

2019-01-29 · Shashank Yadav, Tejas Shimpi, C. Ravindranath Chowdary, Prashant Sharma 외

Segmenting an unordered text document into different sections is a very useful task in many text processing applications like multiple document summarization, question answering, etc. This paper proposes structuring of a…

Document SummarizationQuestion AnsweringSentence

From Faithfulness to Correctness: Generative Reward Models that Think Critically

2025-09-29 · Qiyao Ma, Yunsheng Shi, Hongtao Tian, Chao Wang 외 arxiv

Through reinforcement learning with verifiable rewards (RLVR), large language models have achieved substantial progress in domains with easily verifiable outcomes, such as mathematics and coding. However, when applied to…

Open-Domain Question AnsweringReinforcement Learning

Corpus for Automatic Structuring of Legal Documents

2022-01-31 · LREC 2022 6 · Prathamesh Kalamkar, Aman Tiwari, Astha Agarwal, Saurabh Karn 외

In populous countries, pending legal cases have been growing exponentially. There is a need for developing techniques for processing and organizing legal documents. In this paper, we introduce a new corpus for structurin…

Stylizing ViT: Anatomy-Preserving Instance Style Transfer for Domain Generalization

2026-01-24 · Sebastian Doerrich, Francesco Di Salvo, Jonas Alle, Christian Ledig arxiv

Deep learning models in medical image analysis often struggle with generalizability across domains and demographic groups due to data heterogeneity and scarcity. Traditional augmentation improves robustness, but fails un…

Domain GeneralizationImage ClassificationData AugmentationStyle Transfer