paper-with-me

홈 › Papers

AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation

2025-04-10 · Tuhin Chakrabarty, Philippe Laban, Chien-Sheng Wu

AI-generated text is proliferating across domains, from creative writing and journalism to marketing content and scientific articles. Models can follow user-provided instructions to generate coherent and grammatically correct outputs but in this work, we study a more fundamental question: how do we evaluate and improve the writing quality of AI-generated text? Writing quality assessment has received less attention from the community, in part because it is fundamentally subjective and requires expertise. We first introduce the Writing Quality Benchmark (WQ) by consolidating five writing-preference datasets into 4,729 writing quality judgments. Our experiments show that most of the competitive baselines, including state-of-the-art LLMs that excel at reasoning tasks, barely outperform random baselines on WQ. We then train specialized Writing Quality Reward Models (WQRM) of various sizes for writing quality assessment that demonstrate strong generalization on four out-of-distribution test sets and 74% accuracy on the WQ benchmark. To further show WQRM's practical benefits during inference, we leverage additional test-time compute to generate and rank multiple candidate revisions, allowing us to select higher-quality outputs from an initial draft. Human evaluation with 9 experienced writers confirm that WQRM-based selection produces writing samples preferred by experts 66% overall, and 72.2% when the reward gap is larger than 1 point. We release our datasets and models to encourage community engagement with writing quality assessment and development of AI writing systems better aligned with human preferences.

📄 PDF Abstract BibTeX arXiv:2504.07532

Code (1)

salesforce/creativity_eval 공식 구현 pytorch

Tasks

ArticlesMarketing

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

UDS--DFKI Submission to the WMT2019 Czech--Polish Similar Language Translation Shared Task

2019-08-01 · WS 2019 8 · Santanu Pal, Marcos Zampieri, Josef van Genabith

In this paper we present the UDS-DFKI system submitted to the Similar Language Translation shared task at WMT 2019. The first edition of this shared task featured data from three pairs of similar languages: Czech and Pol…

Translation

Testing Zipf’s meaning-frequency law with wordnets as sense inventories

2019-07-01 · GWC 2019 7 · Francis Bond, Arkadiusz Janz, Marek Maziarz, Ewa Rudnicka

According to George K. Zipf, more frequent words have more senses. We have tested this law using corpora and wordnets of English, Spanish, Portuguese, French, Polish, Japanese, Indonesian and Chinese. We have proved that…

LEMMA

UDS--DFKI Submission to the WMT2019 Similar Language Translation Shared Task

2019-08-16 · Santanu Pal, Marcos Zampieri, Josef van Genabith

In this paper we present the UDS-DFKI system submitted to the Similar Language Translation shared task at WMT 2019. The first edition of this shared task featured data from three pairs of similar languages: Czech and Pol…

Translation

The Polish Sejm Corpus

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk

This document presents the first edition of the Polish Sejm Corpus -- a new specialized resource containing transcribed, automatically annotated utterances of the Members of Polish Sejm (lower chamber of the Polish Parli…

SentenceWord Sense Disambiguation

Evaluation of basic modules for isolated spelling error correction in Polish texts

2019-05-26 · Szymon Rutkowski

Spelling error correction is an important problem in natural language processing, as a prerequisite for good performance in downstream tasks as well as an important feature in user-facing applications. For texts in Polis…