paper-with-me

Papers

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning

2025-08-13 · Yunxiao Wang, Meng Liu, Kaiyu Jiang, Bin Wen, Fan Yang, Tingting Gao, Lizi Liao arxiv

Emotional support conversations require more than fluent responses. Supporters need to understand the seeker's situation and emotions, adopt an appropriate strategy, and respond in a natural, human-like manner. Despite advances in large language models, current systems often lack structured, psychology-informed reasoning. Additionally, it is challenging to enhance these systems through reinforcement learning because of unreliable reward signals. Moreover, reinforcement fine-tuning can amplify repetitive response patterns. We propose structured empathetic reasoning, which breaks support into three steps: conversation history analysis, multimodal emotional state inference, and strategy selection, prior to generating the final reply. To implement this, we introduce SER, a fine-grained dataset with step-level correctness labels and pairwise response preferences. We then present PEER, which uses GRPO with UnifiReward, a unified process-outcome reward model for evaluating both reasoning steps and final responses in multi-turn interactions. To reduce repetition, we enhance data with personality-based rewriting and down-weight redundant outputs. Comprehensive experiments show improved empathy, strategy alignment, and human-likeness without sacrificing diversity. Code and data are available at https://github.com/Yunxiao-Wang/PEER.

📄 PDF Abstract BibTeX arXiv:2508.09521

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Recidivism and Peer Influence with LLM Text Embeddings in Low Security Correctional Facilities

2025-09-25 · Shanjukta Nath, Jiwon Hong, Jae Ho Chang, Keith Warren 외 arxiv

Studying peer effects in language is critical because they often reflect behavioral and personality traits that are important determinants of economic outcomes. However, language is unstructured, non-numeric, and high-di…

NLPeer: A Unified Resource for the Computational Study of Peer Review

2022-11-12 · Nils Dycke, Ilia Kuznetsov, Iryna Gurevych

Peer review constitutes a core component of scholarly publishing; yet it demands substantial expertise and training, and is susceptible to errors and biases. Various applications of NLP for peer reviewing assistance aim …

Scaffolding Collaborative Learning in STEM: A Two-Year Evaluation of a Tool-Integrated Project-Based Methodology

2025-09-02 · Caterina Fuster-Barcelo, Gonzalo R. Rios-Munoz, Arrate Munoz-Barrutia arxiv

This study examines the integration of digital collaborative tools and structured peer evaluation in the Machine Learning for Health master's program, through the redesign of a Biomedical Image Processing course over two…

Quantile Peer Effect Models

2025-06-15 · Aristide Houndetoungan

I propose a flexible structural model to estimate peer effects across various quantiles of the peer outcome distribution. The model allows peers with low, intermediate, and high outcomes to exert distinct influences, the…

Process-Verified Reinforcement Learning for Theorem Proving via Lean

2026-06-18 · Minsu Kim, Se-Young Yun arxiv

While reinforcement learning from verifiable rewards (RLVR) typically has relied on a single binary verification signal, symbolic proof assistants in formal reasoning offer rich, fine-grained structured feedback. This ga…

Reinforcement Learning