paper-with-me

홈 › Papers

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

2026-06-03 · Viktor Veselý, Aleksandar Todorov, Erwan Escudie, Matthia Sabatelli arxiv

Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is poorly understood. We identify a systematic failure mode in deep reinforcement learning (RL) termed Trace-Mediated Peak Bias (TMPB). At intermediate eligibility trace depths, agents irrationally prefer trajectories with high-magnitude reward `peaks'' over alternatives with higher cumulative returns. This provides a mechanistic account of the Peak-End Rule: a human memory bias where experiences are judged by their most intense moments rather than integrated utility. We show that TMPB emerges because traces amplify distal Temporal Difference errors into `gradient shocks'' that fixed-step-size Stochastic Gradient Descent cannot normalize, leading to global overestimation. Conversely, adaptive optimizers mitigate this pathology via second-moment normalization. Our results suggest that human-like saliency distortions may emerge naturally from the mathematical constraints of credit assignment in distributed systems, and that adaptive optimization is a theoretical necessity for rational value estimation.

📄 PDF Abstract BibTeX arXiv:2606.04735

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Bridging Human Interpretation and Machine Representation: A Landscape of Qualitative Data Analysis in the LLM Era

2026-01-16 · Xinyu Pi, Qisen Yang, Chuong Nguyen, Hua Shen arxiv

LLMs are increasingly used to support qualitative research, yet existing systems produce outputs that vary widely--from trace-faithful summaries to theory-mediated explanations and system models. To make these difference…

Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning

2025-07-07 · Jaedong Hwang, Kumar Tanmay, Seok-Jin Lee, Ayush Agrawal 외

Large Language Models (LLMs) have achieved strong performance in domains like mathematics, factual QA, and code generation, yet their multilingual reasoning capabilities in these tasks remain underdeveloped. Especially f…

Code Generationreinforcement-learningReinforcement Learning

Novel Catchbond mediated oscillations in motor-microtubule complexes

2020-05-10 · Sougata Guha, Mithun K. Mitra, Ignacio Pagonabarraga, Sudipto Muhuri

Generation of mechanical oscillation is ubiquitous to wide variety of intracellular processes. We show that catchbonding behaviour of motor proteins provides a generic mechanism of generating spontaneous oscillations in …

Temporal Image Forensics: A Review and Critical Evaluation

2025-09-09 · Robert Jöchl, Andreas Uhl arxiv

Temporal image forensics is the science of estimating the age of a digital image. Usually, time-dependent traces (age traces) introduced by the image acquisition pipeline are exploited for this purpose. In this review, a…

Distilling Information Reliability and Source Trustworthiness from Digital Traces

2016-10-24 · Behzad Tabibian, Isabel Valera, Mehrdad Farajtabar, Le Song 외

Online knowledge repositories typically rely on their users or dedicated editors to evaluate the reliability of their content. These evaluations can be viewed as noisy measurements of both information reliability and inf…