paper-with-me

홈 › Papers

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

2026-06-25 · Han-yu Wang arxiv

Large reasoning models (LRMs) take longer on harder problems, just as humans do, but that surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong it spends more tokens than when it gets that same problem right; humans do the reverse. We separate two levels of deliberation: how response time tracks difficulty across items (registration), and, with item identity fixed, whether an agent spends more on its own failures or successes (allocation). On a public matched human-LRM corpus, thinking LRMs reproduce the known cross-item alignment with human reaction time but diverge from humans within items: on H-ARC every model lands on the opposite side of zero, the four well-powered ones at Cohen's d = 1.47 to 3.13 against -0.10 for humans. Each agent is scored on its own scale; seconds and tokens never share an axis. The dissociation survives item fixed effects and replicates across datasets; a non-thinking baseline shows a wrong-trial expansion of its own but no cross-item alignment, so part of the effect is not specific to reasoning training. We read the human pattern as engagement versus abandonment: people stay on items they expect to solve and give up on the rest. We read the LRM pattern as length driven by uncertainty: chains grow when the model is unsure, exactly when it tends to fail. Under resource-rational metareasoning these are stopping policies that share a difficulty signal but implement opposite control; trace length captures the signal and misses the control.

📄 PDF Abstract BibTeX arXiv:2606.26502

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HERALD: An Annotation Efficient Method to Detect User Disengagement in Social Conversations

2021-06-01 · ACL 2021 5 · Weixin Liang, Kai-Hui Liang, Zhou Yu

Open-domain dialog systems have a user-centric goal: to provide humans with an engaging conversation experience. User engagement is one of the most important metrics for evaluating open-domain dialog systems, and could a…

DenoisingOpen-Domain Dialog

AI Companions as Hyper Attachment and Caregiving Targets

2026-05-15 · Julian De Freitas arxiv

How should we make sense of people's interactions with AI companions-conversational systems built for ongoing, emotionally meaningful relationships? First, I argue these interactions should be understood as attachment re…

Alvorada-Bench: Can Language Models Solve Brazilian University Entrance Exams?

2025-08-19 · Henrique Godoy arxiv

Language models are increasingly used in Brazil, but most evaluation remains English-centric. This paper presents Alvorada-Bench, a 4,515-question, text-only benchmark drawn from five Brazilian university entrance examin…

DRARL: Disengagement-Reason-Augmented Reinforcement Learning for Efficient Improvement of Autonomous Driving Policy

2025-06-20 · Weitao Zhou, Bo Zhang, Zhong Cao, Xiang Li 외

With the increasing presence of automated vehicles on open roads under driver supervision, disengagement cases are becoming more prevalent. While some data-driven planning systems attempt to directly utilize these diseng…

Autonomous DrivingState Estimation

Layered Mutability: Continuity and Governance in Persistent Self-Modifying Agents

2026-04-16 · Krti Tallam arxiv

Persistent language-model agents increasingly combine tool use, tiered memory, reflective prompting, and runtime adaptation. In such systems, behavior is shaped not only by current prompts but by mutable internal conditi…