paper-with-me

홈 › Papers

HorizonBench: Long-Horizon Personalization with Evolving Preferences

2026-04-19 · Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar, Zhongyao Ma, Gelin Zhou, Lin Guan, Na Zhang, Sem Park, Lin Chen, Diyi Yang, Yulia Tsvetkov, Asli Celikyilmaz arxiv

User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent life event. We define this problem as long-horizon personalization and observe that progress on it is limited by data availability and measurement, with no existing resource providing both naturalistic long-horizon interactions and the ground-truth provenance needed to diagnose why models fail. We introduce a data generator that produces conversations from a structured mental state graph, yielding ground-truth provenance for every preference change across 6-month timelines, and from it construct HorizonBench, a benchmark of 4,245 items from 360 simulated users with 6-month conversation histories averaging ~4,300 turns and ~163K tokens. HorizonBench provides a testbed for long-context modeling, memory-augmented architectures, theory-of-mind reasoning, and user modeling. Across 25 frontier models, the best model reaches 52.8% and most score at or below the 20% chance baseline. When these models err on evolved preferences, over a third of the time they select the user's originally stated value without tracking the updated user state. This belief-update failure persists across context lengths and expression explicitness levels, identifying state-tracking capability as the primary bottleneck for long-horizon personalization.

📄 PDF Abstract BibTeX arXiv:2604.17283

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TSUBASA: Improving Long-Horizon Personalization via Evolving Memory and Self-Learning with Context Distillation

2026-04-09 · Xinliang Frederick Zhang, Lu Wang arxiv

Personalized large language models (PLLMs) have garnered significant attention for their ability to align outputs with individual's needs and preferences. However, they still struggle with long-horizon tasks, such as tra…

Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory

2026-05-01 · Derong Xu, Shuochen Liu, Pengfei Luo, Pengyue Jia 외 arxiv

Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving preferences over long interactions. Existing memory systems mainly rely…

Reinforcement Learning

PersonaVLM: Long-Term Personalized Multimodal LLMs

2026-03-20 · Chang Nie, Chaoyou Fu, Yifan Zhang, Haihua Yang 외 arxiv

Multimodal Large Language Models (MLLMs) serve as daily assistants for millions. However, their ability to generate responses aligned with individual preferences remains limited. Prior approaches enable only static, sing…

PRIME: Large Language Model Personalization with Cognitive Memory and Thought Processes

2025-07-07 · Xinliang Frederick Zhang, Nick Beauchamp, Lu Wang

Large language model (LLM) personalization aims to align model outputs with individuals' unique preferences and opinions. While recent efforts have implemented various personalization methods, a unified theoretical frame…

Language ModelingLanguage ModellingLarge Language Model

Towards Natural Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions

2026-03-04 · Qianyun Guo, Yibo Li, Yue Liu, Bryan Hooi arxiv

Large Language Models (LLMs) are increasingly serving as personal assistants, where users may share individual preferences over extended interactions. However, assessing how well LLMs can follow these preferences in natu…