paper-with-me

홈 › Papers

Behavior-Aware Auxiliary Corrections for Off-Policy Temporal-Difference Prediction

2026-05-17 · Xingguo Chen, Zhiang He, Yuchen Shen, Shangdong Yang, Chao Li, Guang Yang, Wenhao Wang arxiv

Temporal-difference learning with function approximation can be unstable under off-policy sampling. TDC stabilizes off-policy TD through an auxiliary covariance correction, and TDRC further regularizes this correction in a single-timescale recursion. This paper studies a behavior-aware replacement of the auxiliary covariance geometry in the linear prediction setting, which is the standard local model for understanding the feature-space dynamics of value-function approximation. We first replace the TDC auxiliary matrix (C) by the behavior Bellman matrix (A_μ), yielding BA-TDC, and then regularize the same behavior-aware equation to obtain BA-TDRC. This two-step construction separates the contribution of behavior-aware geometry from the contribution of regularization. The linear analysis also provides a tractable model for an auxiliary-geometry design question that arises in neural-network value approximation, where feature covariances and temporal transition matrices jointly shape the last-layer correction dynamics. We give a finite-state mean-system formulation, prove fixed-point preservation and almost-sure convergence under a Hurwitz stability condition on the instantiated mean system, and compare deterministic mean rates through the spectral radius of the exact linear error recursion. Experiments on the two-state counterexample, Baird's counterexample, Random Walk, and Boyan Chain show that the behavior-aware replacement can be highly beneficial by itself on some tasks, but that regularization is necessary for robust performance across harder settings.

📄 PDF Abstract BibTeX arXiv:2605.28855

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Q($λ$) with Off-Policy Corrections

2016-02-16 · Anna Harutyunyan, Marc G. Bellemare, Tom Stepleton, Remi Munos

We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target p…

Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction

2026-05-16 · Xingguo Chen, Yuchen Shen, Shangdong Yang, Chao Li 외 arxiv

Gradient temporal-difference methods provide stable off-policy prediction with linear function approximation, but their practical performance is strongly affected by the geometry induced by the auxiliary-variable metric.…

Set-Supervised Diffusion Policy: Learning Action-Chunking Diffusion through Corrections

2026-06-01 · Zhaoting Li, Gang Chen, Javier Alonso-Mora, Cosimo Della Santina 외 arxiv

Diffusion policies have recently emerged as a powerful framework for robotic manipulation. However, like other behavior cloning methods, they remain vulnerable to distributional shift, often requiring human-in-the-loop i…

STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning

2026-06-29 · Zhihao Liu, Qiuyi Gu, Yitao Wang, Dongming Qiao 외 arxiv

Real-world robot learning increasingly relies on heterogeneous data, but demonstrations and rollouts often mix useful progress with stalls, corrections, and suboptimal behavior. Effective policy learning therefore requir…

Mixture Policy based Multi-Hop Reasoning over N-tuple Temporal Knowledge Graphs

2025-05-19 · Zhongni Hou, Miao Su, Xiaolong Jin, Zixuan Li 외

Temporal Knowledge Graphs (TKGs), which utilize quadruples in the form of (subject, predicate, object, timestamp) to describe temporal facts, have attracted extensive attention. N-tuple TKGs (N-TKGs) further extend tradi…

Knowledge Graphs