The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
Standard offline evaluations for language models -- a series of independent, state-less inferences made by models -- fail to capture how language models actually behave in practice, where personalization fundamentally alters model behavior. For instance, identical benchmark questions to the same language model can produce markedly different responses when prompted to a state-less system, in one user's chat session, or in a different user's chat session. In this work, we provide empirical evidence showcasing this phenomenon by comparing offline evaluations to field evaluations conducted by having 800 real users of ChatGPT and Gemini pose benchmark and other provided questions to their chat interfaces.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Unveiling Bias in Fairness Evaluations of Large Language Models: A Critical Literature Review of Music and Movie Recommendation Systems
The rise of generative artificial intelligence, particularly Large Language Models (LLMs), has intensified the imperative to scrutinize fairness alongside accuracy. Recent studies have begun to investigate fairness evalu…
FairnessMovie RecommendationRecommendation SystemsLearning to Rank in the Position Based Model with Bandit Feedback
Personalization is a crucial aspect of many online experiences. In particular, content ranking is often a key component in delivering sophisticated personalization results. Commonly, supervised learning-to-rank methods a…
Learning-To-RankMulti-Armed BanditsPositionThompson SamplingLanguage Model Personalization via Reward Factorization
Modern large language models (LLMs) are optimized for human-aligned responses using Reinforcement Learning from Human Feedback (RLHF). However, existing RLHF approaches assume a universal preference model and fail to acc…
Language ModelingLanguage ModellingmodelEnd-to-end Offline Reinforcement Learning for Glycemia Control
The development of closed-loop systems for glycemia control in type I diabetes relies heavily on simulated patients. Improving the performances and adaptability of these close-loops raises the risk of over-fitting the si…
Offline RLreinforcement-learningReinforcement LearningLived Experience in Dialogue: Co-designing Personalization in Large Language Models to Support Youth Mental Well-being
Youth increasingly turn to large language models (LLMs) for mental well-being support, yet current personalization in LLMs can overlook the heterogeneous lived experiences shaping their needs. We conducted a participator…