paper-with-me

Papers

PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs

2025-08-08 · Xiao Fu, Hossein A. Rahmani, Bin Wu, Jerome Ramos, Emine Yilmaz, Aldo Lipani arxiv

Personalised text generation is essential for user-centric information systems, yet most evaluation methods overlook the individuality of users. We introduce \textbf{PREF}, a \textbf{P}ersonalised \textbf{R}eference-free \textbf{E}valuation \textbf{F}ramework that jointly measures general output quality and user-specific alignment without requiring gold personalised references. PREF operates in a three-step pipeline: (1) a coverage stage uses a large language model (LLM) to generate a comprehensive, query-specific guideline covering universal criteria such as factuality, coherence, and completeness; (2) a preference stage re-ranks and selectively augments these factors using the target user's profile, stated or inferred preferences, and context, producing a personalised evaluation rubric; and (3) a scoring stage applies an LLM judge to rate candidate answers against this rubric, ensuring baseline adequacy while capturing subjective priorities. This separation of coverage from preference improves robustness, transparency, and reusability, and allows smaller models to approximate the personalised quality of larger ones. Experiments on the PrefEval benchmark, including implicit preference-following tasks, show that PREF achieves higher accuracy, better calibration, and closer alignment with human judgments than strong baselines. By enabling scalable, interpretable, and user-aligned evaluation, PREF lays the groundwork for more reliable assessment and development of personalised language generation systems.

📄 PDF Abstract BibTeX arXiv:2508.10028

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users

2026-05-13 · Hannah Rose Kirk, Liu Leqi, Fanzhi Zeng, Henry Davidson 외 arxiv

Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simulated users rather than real people. Thi…

Modelling User Preferences using Word Embeddings for Context-Aware Venue Recommendation

2016-06-24 · Manotumruksa Jarana, Macdonald Craig, Ounis Iadh

Venue recommendation aims to assist users by making personalised suggestions of venues to visit, building upon data available from location-based social networks (LBSNs) such as Foursquare. A particular challenge for thi…

Word Embeddings

Trainable Referring Expression Generation using Overspecification Preferences

2017-04-12 · Thiago castro Ferreira, Ivandre Paraboni

Referring expression generation (REG) models that use speaker-dependent information require a considerable amount of training data produced by every individual speaker, or may otherwise perform poorly. In this work we pr…

Referring ExpressionReferring expression generation

A Personalised Formal Verification Framework for Monitoring Activities of Daily Living of Older Adults Living Independently in Their Homes

2025-07-11 · Ricardo Contreras, Filip Smola, Nuša Farič, Jiawei Zheng 외 arxiv

There is an imperative need to provide quality of life to a growing population of older adults living independently. Personalised solutions that focus on the person and take into consideration their preferences and conte…

Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback

2023-03-09 · Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, Scott A. Hale

Large language models (LLMs) are used to generate content for a wide range of tasks, and are set to reach a growing audience in coming years due to integration in product interfaces like ChatGPT or search engines like Bi…

Red Teaming