paper-with-me

홈 › Papers

UserLM-R1: Modeling Human Reasoning in User Language Models with Multi-Reward Reinforcement Learning

2026-01-14 · Feng Zhang, Shijia Li, Chunmao Zhang, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, Jingwen Xu, Han Liu arxiv

User simulators serve as the critical interactive environment for agent post-training, and an ideal user simulator generalizes across domains and proactively engages in negotiation by challenging or bargaining. However, current methods exhibit two issues. They rely on static and context-unaware profiles, necessitating extensive manual redesign for new scenarios, thus limiting generalizability. Moreover, they neglect human strategic thinking, leading to vulnerability to agent manipulation. To address these issues, we propose UserLM-R1, a novel user language model with reasoning capability. Specifically, we first construct comprehensive user profiles with both static roles and dynamic scenario-specific goals for adaptation to diverse scenarios. Then, we propose a goal-driven decision-making policy to generate high-quality rationales before producing responses, and further refine the reasoning and improve strategic capabilities with supervised fine-tuning and multi-reward reinforcement learning. Extensive experimental results demonstrate that UserLM-R1 outperforms competitive baselines, particularly on the more challenging adversarial set.

📄 PDF Abstract BibTeX arXiv:2601.09215

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling

2026-05-26 · Alan Zhu, Mihran Miroyan, Carolyn Wang, Andrew Zhou 외 arxiv

User modeling aims to use language models (LMs) to mimic an individual's behavior from a corpus of past context-action pairs (e.g., conversation turns), enabling the simulation of users in settings like behavioral scienc…

Zero-Shot Human Mobility Forecasting via Large Language Model with Hierarchical Reasoning

2025-09-20 · Wenyao Li, Ran Zhang, Pengyang Wang, Yuanchun Zhou 외 arxiv

Human mobility forecasting is important for applications such as transportation planning, urban management, and personalized recommendations. However, existing methods often fail to generalize to unseen users or location…

Question Answering

A Unified Spoken Language Model with Injected Emotional-Attribution Thinking for Human-like Interaction

2026-01-08 · Qing Wang, Zehan Li, Yaodong Song, Hongjie Chen 외 arxiv

This paper presents a unified spoken language model for emotional intelligence, enhanced by a novel data construction strategy termed Injected Emotional-Attribution Thinking (IEAT). IEAT incorporates user emotional state…

Empathetic Response GenerationEmotional IntelligenceTrajectory Modeling

CoAX: Cognitive-Oriented Attribution eXplanation User Model of Human Understanding of AI Explanations

2026-04-30 · Louth Bin Rawshan, Zhuoyu Wang, Brian Y. Lim arxiv

Explainable AI (XAI) aims to improve user understanding and decisions when using AI models. However, despite innovations in XAI, recent user evaluations reveal that this goal remains elusive. Understanding human cognitio…

Feature Importance

IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models

2026-04-27 · Hamed Rahimi, Clemence Grislain, Adrien Jacquet Cretides, Olivier Sigaud 외 arxiv

Improving the effectiveness of human-robot interaction requires social robots to accurately infer human goals through robust intention understanding. This challenge is particularly critical in multimodal settings, where …