paper-with-me

홈 › Papers

Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users

2025-10-20 · Melik Ozolcer, Sang Won Bae arxiv

We study a web-deployed, tool-augmented LLM health coach with real users. In a pilot with seven users (280 rated turns), offline policy evaluation (OPE) over factorized decision heads (Tool/Style) shows that a uniform heavy-tool policy raises average value on logs but harms specific subgroups, most notably low-health-literacy/high-self-efficacy users. A lightweight simulator with hidden archetypes further shows that adding a small early information-gain bonus reliably shortens trait identification and improves goal success and pass@3. Together, these early findings indicate an evaluation-first path to personalization: freeze the generator, learn subgroup-aware decision heads on typed rewards (objective tool outcomes and satisfaction), and always report per-archetype metrics to surface subgroup harms that averages obscure.

📄 PDF Abstract BibTeX arXiv:2510.17173

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators

2024-05-27 · Allen Nie, Yash Chandak, Christina J. Yuan, Anirudhan Badrinath 외

Offline policy evaluation (OPE) allows us to evaluate and estimate a new sequential decision-making policy's performance by leveraging historical interaction data collected from other policies. Evaluating a new policy on…

Decision MakingOffline RLOff-policy evaluationSequential Decision Making

Distributional Offline Policy Evaluation with Predictive Error Guarantees

2023-02-19 · Runzhe Wu, Masatoshi Uehara, Wen Sun

We study the problem of estimating the distribution of the return of a policy using an offline dataset that is not generated from the policy, i.e., distributional offline policy evaluation (OPE). We propose an algorithm …

Avoiding Overfitting to the Importance Weights in Offline Policy Optimization

2021-09-29 · Yao Liu, Emma Brunskill

Offline policy optimization has a critical impact on many real-world decision-making problems, as online learning is costly and concerning in many applications. Importance sampling and its variants are a widely used type…

Decision Making

Offline Policy Optimization with Eligible Actions

2022-07-01 · Yao Liu, Yannis Flet-Berliac, Emma Brunskill

Offline policy optimization could have a large impact on many real-world decision-making problems, as online learning may be infeasible in many applications. Importance sampling and its variants are a commonly used type …

continuous-controlContinuous ControlDecision Making

COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation

2022-04-19 · ICLR 2022 4 · Jongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 외

We consider the offline constrained reinforcement learning (RL) problem, in which the agent aims to compute a policy that maximizes expected return while satisfying given cost constraints, learning only from a pre-collec…

Offline RLOff-policy evaluationreinforcement-learningReinforcement Learning+1