paper-with-me

홈 › Papers

Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures

2025-10-16 · Shuangshuang Ying, Yunwen Li, Xingwei Qu, Xin Li, Sheng Jin, Minghao Liu, Zhoufutu Wen, Xeron Du, Tianyu Zheng, Yichi Zhang, Letian Ni, Yuyang Cheng, Zhenzhu Yang, Qiguang Chen, Jingzhe Ding, Shengda Long, Wangchunshu Zhou, Jiazhan Feng, Wanjun Zhong, Libo Qin, Ge Zhang, Wenhao Huang, Wanxiang Che, Chenghua Lin arxiv

Current preference learning methods achieve high accuracy on standard benchmarks but exhibit significant performance degradation when objective quality signals are removed. We introduce WritingPreferenceBench, a dataset of 1,800 human-annotated preference pairs (1,200 English, 600 Chinese) across 8 creative writing genres, where responses are matched for objective correctness, factual accuracy, and length. On this benchmark, sequence-based reward models--the standard architecture for RLHF--achieve only 52.7% mean accuracy, while zero-shot language model judges perform at 53.9%. In contrast, generative reward models that produce explicit reasoning chains achieve 81.8% accuracy. We observe high within-model variance across genres: individual models range from 18.2% to 81.8% accuracy across different writing categories, with standard deviations averaging 10.1%. This variance persists regardless of model scale, with 27B parameter models showing no consistent improvement over 8B variants. Our results suggest that current RLHF methods primarily learn to detect objective errors rather than capture subjective quality preferences (e.g., creativity, stylistic flair, and emotional resonance), and that successful preference modeling may require intermediate reasoning representations rather than direct classification.

📄 PDF Abstract BibTeX arXiv:2510.14616

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

2025-08-26 · Jianxing Liao, Tian Zhang, Xiao Feng, Yusong Zhang 외 arxiv

Large language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emotional expression) and objective constrai…

Reinforcement LearningInstruction Following

Minimal Macro-Based Rewritings of Formal Languages: Theory and Applications in Ontology Engineering (and beyond)

2023-12-18 · Christian Kindermann, Anne-Marie George, Bijan Parsia, Uli Sattler

In this paper, we introduce the problem of rewriting finite formal languages using syntactic macros such that the rewriting is minimal in size. We present polynomial-time algorithms to solve variants of this problem and …

Controllable Complementarity: Subjective Preferences in Human-AI Collaboration

2025-03-07 · Chase McDonald, Cleotilde Gonzalez

Research on human-AI collaboration often prioritizes objective performance. However, understanding human subjective preferences is essential to improving human-AI complementarity and human experiences. We investigate hum…

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness

2026-05-11 · Ivo Petrov, Jasper Dekoninck, Dimitar I. Dimitrov, Martin Vechev arxiv

Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correctness alone is not sufficient: mathematical proofs should also be cle…

Mathematical Reasoning

DEAR: Dataset for Evaluating the Aesthetics of Rendering

2025-12-04 · Vsevolod Plohotnuk, Artyom Panshin, Nikola Banić, Simone Bianco 외 arxiv

Traditional Image Quality Assessment~(IQA) focuses on quantifying technical degradations such as noise, blur, or compression artifacts, using both full-reference and no-reference objective metrics. However, evaluation of…

Image Quality Assessment