paper-with-me

홈 › Papers

Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement Learning

2025-09-05 · Bastien Dubail, Stefan Stojanovic, Alexandre Proutière arxiv

Low-rank structure is a common implicit assumption in many modern reinforcement learning (RL) algorithms. For instance, reward-free and goal-conditioned RL methods often presume that the successor measure admits a low-rank representation. In this work, we challenge this assumption by first remarking that the successor measure itself is not approximately low-rank. Instead, we demonstrate that a low-rank structure naturally emerges in the shifted successor measure, which captures the system dynamics after bypassing a few initial transitions. We provide finite-sample performance guarantees for the entry-wise estimation of a low-rank approximation of the shifted successor measure from sampled entries. Our analysis reveals that both the approximation and estimation errors are primarily governed by a newly introduced quantitity: the spectral recoverability of the corresponding matrix. To bound this parameter, we derive a new class of functional inequalities for Markov chains that we call Type II Poincaré inequalities and from which we can quantify the amount of shift needed for effective low-rank approximation and estimation. This analysis shows in particular that the required shift depends on decay of the high-order singular values of the shifted successor measure and is hence typically small in practice. Additionally, we establish a connection between the necessary shift and the local mixing properties of the underlying dynamical system, which provides a natural way of selecting the shift. Finally, we validate our theoretical findings with experiments, and demonstrate that shifting the successor measure indeed leads to improved performance in goal-conditioned RL.

📄 PDF Abstract BibTeX arXiv:2509.05193

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Using neural topic models to track context shifts of words: a case study of COVID-related terms before and after the lockdown in April 2020

2022-05-01 · LChange (ACL) 2022 5 · Olga Kellert, Md Mahmud Uz Zaman

This paper explores lexical meaning changes in a new dataset, which includes tweets from before and after the COVID-related lockdown in April 2020. We use this dataset to evaluate traditional and more recent unsupervised…

Language ModelingLanguage ModellingTopic Models

jina-reranker-v3: Last but Not Late Interaction for Listwise Document Reranking

2025-09-29 · Feng Wang, Yuqing Li, Han Xiao arxiv

jina-reranker-v3 is a 0.6B-parameter multilingual listwise reranker that introduces a novel "last but not late" interaction. Unlike late interaction models like ColBERT that encode documents separately before multi-vecto…

Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN

2025-05-22 · Yao Xu, Mingyu Xu, Fangyu Lei, Wangtao Sun 외

Recently, models such as OpenAI-o1 and DeepSeek-R1 have demonstrated remarkable performance on complex reasoning tasks through Long Chain-of-Thought (Long-CoT) reasoning. Although distilling this capability into student …

Mathematical Reasoning

Shift-Invariant Attribute Scoring for Kolmogorov-Arnold Networks via Shapley Value

2025-10-02 · Wangxuan Fan, Ching Wang, Siqi Li, Nan Liu arxiv

For many real-world applications, understanding feature-outcome relationships is as crucial as achieving high predictive accuracy. While traditional neural networks excel at prediction, their black-box nature obscures un…

Network Pruning

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning

2026-04-30 · Sahil Mishra, Srinitish Srinivasan, Sourish Dasgupta, Tanmoy Chakraborty arxiv

Real-world knowledge is often organized as hierarchies such as product taxonomies, medical ontologies, and label trees, yet learning hierarchical representations is challenging due to asymmetric structure and noisy seman…