paper-with-me

홈 › Papers

LikeBench: Evaluating Subjective Likability in LLMs for Personalization

2025-12-15 · Md Awsafur Rahman, Adam Gabrys, Doug Kang, Jingjing Sun, Tian Tan, Ashwin Chandramouli arxiv

A personalized LLM should remember user facts, apply them correctly, and adapt over time to provide responses that the user prefers. Existing LLM personalization benchmarks are largely centered on two axes: accurately recalling user information and accurately applying remembered information in downstream tasks. We argue that a third axis, likability, is both subjective and central to user experience, yet under-measured by current benchmarks. To measure likability holistically, we introduce LikeBench, a multi-session, dynamic evaluation framework that measures likability across multiple dimensions by how much an LLM can adapt over time to a user's preferences to provide more likable responses. In LikeBench, the LLMs engage in conversation with a simulated user and learn preferences only from the ongoing dialogue. As the interaction unfolds, models try to adapt to responses, and after each turn, they are evaluated for likability across seven dimensions by the same simulated user. To the best of our knowledge, we are the first to decompose likability into multiple diagnostic metrics: emotional adaptation, formality matching, knowledge adaptation, reference understanding, conversation length fit, humor fit, and callback, which makes it easier to pinpoint where a model falls short. To make the simulated user more realistic and discriminative, LikeBench uses fine-grained, psychologically grounded descriptive personas rather than the coarse high/low trait rating based personas used in prior work. Our benchmark shows that strong memory performance does not guarantee high likability: DeepSeek R1, with lower memory accuracy (86%, 17 facts/profile), outperformed Qwen3 by 28% on likability score despite Qwen3's higher memory accuracy (93%, 43 facts/profile). Even SOTA models like GPT-5 adapt well in short exchanges but show only limited robustness in longer, noisier interactions.

📄 PDF Abstract BibTeX arXiv:2512.13077

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data

2024-07-05 · Hitoshi Suda, Aya Watanabe, Shinnosuke Takamichi

This paper introduces CocoNut-Humoresque, an open-source large-scale speech likability corpus that includes speech segments and their per-listener likability scores. Evaluating voice likability is essential to designing …

Exploring Safety-Utility Trade-Offs in Personalized Language Models

2024-06-17 · Anvesh Rao Vijjini, Somnath Basu Roy Chowdhury, Snigdha Chaturvedi

As large language models (LLMs) become increasingly integrated into daily applications, it is essential to ensure they operate fairly across diverse user demographics. In this work, we show that LLMs suffer from personal…

General Knowledge

A Genre-Aware Attention Model to Improve the Likability Prediction of Books

2018-10-01 · EMNLP 2018 10 · Suraj Maharjan, Manuel Montes, Fabio A. Gonz{\'a}lez, Thamar Solorio

Likability prediction of books has many uses. Readers, writers, as well as the publishing industry, can all benefit from automatic book likability prediction systems. In order to make reliable decisions, these systems ne…

Feature ImportancePrediction

PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization

2025-06-15 · Meiling Tao, Chenghao Zhu, Dongyi Ding, Tiannan Wang 외

With the rapid improvement in the general capabilities of LLMs, LLM personalization, i.e., how to build LLM systems that can generate personalized responses or services that are tailored to distinct user personas, has be…

Personalized Large Language Models

2024-02-14 · Stanisław Woźniak, Bartłomiej Koptyra, Arkadiusz Janz, Przemysław Kazienko 외

Large language models (LLMs) have significantly advanced Natural Language Processing (NLP) tasks in recent years. However, their universal nature poses limitations in scenarios requiring personalized responses, such as r…

Emotion RecognitionHate Speech DetectionRecommendation Systems