paper-with-me

홈 › Papers

Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework

2026-05-15 · Xintong Yao arxiv

Long-term interaction with LLM-based systems may produce alignment drift: a gradual process in which system outputs become less constrained by the user's current message and more shaped by prior interaction history, while still appearing helpful, coherent, and responsive. This process is difficult to detect because the user's subjective experience may improve as the system becomes more familiar, useful, and attuned. Existing research on human-LLM interaction has largely focused on short-term task performance, isolated outputs, or single-instance alignment problems, leaving slow and cumulative interaction-level dynamics undercharacterized. This paper proposes a mechanism-oriented framework for describing alignment drift. The framework defines the distinction between signal A and signal B, explains how drift develops through feedback loops and sub-pattern selection, divides the process into three interactional regimes, and identifies boundary conditions for controlling drift. By framing alignment drift as a recursive interactional process rather than an isolated model-side failure, the paper provides a conceptual basis for studying long-term human-system interaction.

📄 PDF Abstract BibTeX arXiv:2605.16516

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Alignment Drift in CEFR-prompted LLMs for Interactive Spanish Tutoring

2025-05-13 · Mina Almasi, Ross Deans Kristensen-McLachlan

This paper investigates the potentials of Large Language Models (LLMs) as adaptive tutors in the context of second-language learning. In particular, we evaluate whether system prompting can reliably constrain LLMs to gen…

BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents

2026-03-25 · Praveen Kumar Myakala, Manan Agrawal, Rahul Manche arxiv

LLMs are increasingly used as long-running conversational agents, yet every major benchmark evaluating their memory treats user information as static facts to be stored and retrieved. That's the wrong model. People chang…

Improving Long-Term Metrics in Recommendation Systems using Short-Horizon Reinforcement Learning

2021-06-01 · Bogdan Mazoure, Paul Mineiro, Pavithra Srinath, Reza Sharifi Sedeh 외

We study session-based recommendation scenarios where we want to recommend items to users during sequential interactions to improve their long-term utility. Optimizing a long-term metric is challenging because the learni…

Offline RLRecommendation Systemsreinforcement-learningReinforcement Learning (RL)+1

Drift: Decoding-time Personalized Alignments with Implicit User Preferences

2025-02-20 · Minbeom Kim, Kang-il Lee, Seongho Joo, Hwaran Lee 외

Personalized alignments for individual users have been a long-standing goal in large language models (LLMs). We introduce Drift, a novel framework that personalizes LLMs at decoding time with implicit user preferences. T…

Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety

2025-10-11 · Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao 외 arxiv

As large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards ena…