paper-with-me

홈 › Papers

Attention Limited Reward Learning

2026-07-06 · Wenqian Xing arxiv

Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences. RLHF and related alignment pipelines typically model such comparisons with Bradley--Terry log-odds, where choice probabilities are governed by latent reward differences. This paper examines what this assumption misses through a reduced-form model motivated by rational inattention, in which each label is generated by a low-capacity evaluation channel. The model separates two forms of ambiguity that standard reward modeling tends to conflate: a comparison may be difficult because the two candidates are genuinely close in value, or because the relevant distinction is hard to detect under limited attention. We show that limited attention can fundamentally distort what pairwise comparisons reveal. In particular, passive comparison data cannot generally distinguish reward, attention, and default tendencies, and heterogeneous attention can make standard Bradley--Terry reward modeling recover misleading rankings. Our analysis shows that learning is governed not by the raw number of labels, but by the amount of attended information each label carries. A case study on human votes over language-model pairs from Chatbot Arena exhibits the predicted signature, a cyclic component of the comparison data that exceeds sampling noise and that no scalar reward can represent; a second case study on perceptual comparisons shows that response times and gaze carry gap information that the labels do not. This perspective suggests that human feedback should be treated not as direct revealed preference, but as an attention-limited measurement process: a weak preference signal may reflect hidden evaluation difficulty rather than genuine indifference.

📄 PDF Abstract BibTeX arXiv:2607.04590

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Attentional Model of Time Discounting

2025-05-19 · Zijian Zark Wang

When decision makers evaluate a sequence of rewards, they may pay more attention to larger rewards and, given attention is limited, less attention to smaller rewards. They may also become less attentive to each reward wh…

model

Dealing with Sparse Rewards Using Graph Neural Networks

2022-03-25 · Matvey Gerasyov, Ilya Makarov

Deep reinforcement learning in partially observable environments is a difficult task in itself, and can be further complicated by a sparse reward signal. Most tasks involving navigation in three-dimensional environments …

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Attention-Based Reward Shaping for Sparse and Delayed Rewards

2025-05-16 · Ian Holmes, Min Chi

Sparse and delayed reward functions pose a significant obstacle for real-world Reinforcement Learning (RL) applications. In this work, we propose Attention-based REward Shaping (ARES), a general and robust algorithm whic…

Reinforcement Learning (RL)

Enhancing Code LLM Training with Programmer Attention

2025-03-19 · Yifan Zhang, Chen Huang, Zachary Karas, Dung Thuy Nguyen 외

Human attention provides valuable yet underexploited signals for code LLM training, offering a perspective beyond purely machine-driven attention. Despite the complexity and cost of collecting eye-tracking data, there ha…

Code Summarization

Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs

2026-02-09 · Siqu Ou, Tianrui Wan, Zhiyuan Zhao, Junyu Gao 외 arxiv

While chain-of-thought (CoT) reasoning has substantially improved multimodal large language models (MLLMs) on complex reasoning tasks, existing approaches largely rely on long textual reasoning trajectories and provide l…

Reinforcement LearningVisual Reasoning