paper-with-me

홈 › Papers

The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior

2025-09-18 · Angelina Wang, Daniel E. Ho, Sanmi Koyejo arxiv

Standard offline evaluations for language models -- a series of independent, state-less inferences made by models -- fail to capture how language models actually behave in practice, where personalization fundamentally alters model behavior. For instance, identical benchmark questions to the same language model can produce markedly different responses when prompted to a state-less system, in one user's chat session, or in a different user's chat session. In this work, we provide empirical evidence showcasing this phenomenon by comparing offline evaluations to field evaluations conducted by having 800 real users of ChatGPT and Gemini pose benchmark and other provided questions to their chat interfaces.

📄 PDF Abstract BibTeX arXiv:2509.19364

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unveiling Bias in Fairness Evaluations of Large Language Models: A Critical Literature Review of Music and Movie Recommendation Systems

2024-01-08 · Chandan Kumar Sah, Dr. Lian Xiaoli, Muhammad Mirajul Islam

The rise of generative artificial intelligence, particularly Large Language Models (LLMs), has intensified the imperative to scrutinize fairness alongside accuracy. Recent studies have begun to investigate fairness evalu…

FairnessMovie RecommendationRecommendation Systems

Learning to Rank in the Position Based Model with Bandit Feedback

2020-04-27 · Beyza Ermis, Patrick Ernst, Yannik Stein, Giovanni Zappella

Personalization is a crucial aspect of many online experiences. In particular, content ranking is often a key component in delivering sophisticated personalization results. Commonly, supervised learning-to-rank methods a…

Learning-To-RankMulti-Armed BanditsPositionThompson Sampling

Language Model Personalization via Reward Factorization

2025-03-08 · Idan Shenfeld, Felix Faltings, Pulkit Agrawal, Aldo Pacchiano

Modern large language models (LLMs) are optimized for human-aligned responses using Reinforcement Learning from Human Feedback (RLHF). However, existing RLHF approaches assume a universal preference model and fail to acc…

Language ModelingLanguage Modellingmodel

End-to-end Offline Reinforcement Learning for Glycemia Control

2023-10-16 · Tristan Beolet, Alice Adenis, Erik Huneker, Maxime Louis

The development of closed-loop systems for glycemia control in type I diabetes relies heavily on simulated patients. Improving the performances and adaptability of these close-loops raises the risk of over-fitting the si…

Offline RLreinforcement-learningReinforcement Learning

Lived Experience in Dialogue: Co-designing Personalization in Large Language Models to Support Youth Mental Well-being

2025-11-07 · Kathleen W. Guan, Sarthak Giri, Mohammed Amara, Bernard J. Jansen 외 arxiv

Youth increasingly turn to large language models (LLMs) for mental well-being support, yet current personalization in LLMs can overlook the heterogeneous lived experiences shaping their needs. We conducted a participator…