paper-with-me

홈 › Papers

PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization

2025-06-15 · Meiling Tao, Chenghao Zhu, Dongyi Ding, Tiannan Wang, Yuchen Eleanor Jiang, Wangchunshu Zhou

With the rapid improvement in the general capabilities of LLMs, LLM personalization, i.e., how to build LLM systems that can generate personalized responses or services that are tailored to distinct user personas, has become an increasingly important research and engineering problem. However, unlike many new challenging benchmarks being released for evaluating the general/reasoning capabilities, the lack of high-quality benchmarks for evaluating LLM personalization greatly hinders progress in this field. To address this, we introduce PersonaFeedback, a new benchmark that directly evaluates LLMs' ability to provide personalized responses given pre-defined user personas and queries. Unlike existing benchmarks that require models to infer implicit user personas from historical interactions, PersonaFeedback decouples persona inference from personalization, focusing on evaluating the model's ability to generate responses tailored to explicit personas. PersonaFeedback consists of 8298 human-annotated test cases, which are categorized into easy, medium, and hard tiers based on the contextual complexity of the user personas and the difficulty in distinguishing subtle differences between two personalized responses. We conduct comprehensive evaluations across a wide range of models. The empirical results reveal that even state-of-the-art LLMs that can solve complex real-world reasoning tasks could fall short on the hard tier of PersonaFeedback where even human evaluators may find the distinctions challenging. Furthermore, we conduct an in-depth analysis of failure modes across various types of systems, demonstrating that the current retrieval-augmented framework should not be seen as a de facto solution for personalization tasks. All benchmark data, annotation protocols, and the evaluation pipeline will be publicly available to facilitate future research on LLM personalization.

📄 PDF Abstract BibTeX arXiv:2506.12915

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch

2025-02-24 · Xueru Wen, Jie Lou, Zichao Li, Yaojie Lu 외

Reward models (RMs) are crucial for aligning large language models (LLMs) with human preferences. However, most RM research is centered on English and relies heavily on synthetic resources, which leads to limited and les…

PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding

2017-03-22 · Chunhui Liu, Yueyu Hu, Yanghao Li, Sijie Song 외

Despite the fact that many 3D human activity benchmarks being proposed, most existing action datasets focus on the action recognition tasks for the segmented videos. There is a lack of standard large-scale benchmarks, es…

Action DetectionAction RecognitionAction UnderstandingTemporal Action Localization

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

2026-07-08 · Sirui Zhang, Tianle Wang, Xinyi Tong, Peiyang Yu 외 arxiv

Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual judgments. Progress in this area has been limited by the lack of large-s…

CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents

2026-03-04 · Martin Kostelník, Michal Hradiš, Martin Dočekal arxiv

Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based on Czech historical documents, containing…

Humans in Kitchens: A Dataset for Multi-Person Human Motion Forecasting with Scene Context

2023-09-26 · NeurIPS 2023 11

Forecasting human motion of multiple persons is very challenging. It requires to model the interactions between humans and the interactions with objects and the environment. For example, a person might want to make a cof…