paper-with-me

홈 › Papers

Learning values across many orders of magnitude

2016-02-24 · NeurIPS 2016 12 · Hado van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih, David Silver

Most learning algorithms are not invariant to the scale of the function that is being approximated. We propose to adaptively normalize the targets used in learning. This is useful in value-based reinforcement learning, where the magnitude of appropriate value approximations can change over time when we update the policy of behavior. Our main motivation is prior work on learning to play Atari games, where the rewards were all clipped to a predetermined range. This clipping facilitates learning across many different games with a single learning algorithm, but a clipped reward function can result in qualitatively different behavior. Using the adaptive normalization we can remove this domain-specific heuristic without diminishing overall performance.

📄 PDF Abstract BibTeX arXiv:1602.07714

Code (0)

등록된 구현이 없습니다.

Tasks

Atari Gamesreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

PDD-SHAP: Fast Approximations for Shapley Values using Functional Decomposition

2022-08-26 · Arne Gevaert, Yvan Saeys

Because of their strong theoretical properties, Shapley values have become very popular as a way to explain predictions made by black box models. Unfortuately, most existing techniques to compute Shapley values are compu…

FastSHAP: Real-Time Shapley Value Estimation

2021-07-15 · ICLR 2022 4 · Neil Jethani, Mukund Sudarshan, Ian Covert, Su-In Lee 외

Shapley values are widely used to explain black-box models, but they are costly to calculate because they require many model evaluations. We introduce FastSHAP, a method for estimating Shapley values in a single forward …

Exactly Computing do-Shapley Values

2026-02-06 · R. Teal Witter, Álvaro Parafita, Tomas Garriga, Maximilian Muschalik 외 arxiv

Structural Causal Models (SCM) are a powerful framework for describing complicated dynamics across the natural sciences. A particularly elegant way of interpreting SCMs is do-Shapley, a game-theoretic method of quantifyi…

Only relative ranks matter in weight-clustered large language models

2026-03-18 · Borja Aizpurua, Sukhbinder Singh, Román Orús arxiv

Large language models (LLMs) contain billions of parameters, yet many exact values are not essential. We show that what matters most is the relative rank of weights-whether one connection is stronger or weaker than anoth…

Model Compression

Variance Reduction in Monte Carlo Counterfactual Regret Minimization (VR-MCCFR) for Extensive Form Games using Baselines

2018-09-09 · Martin Schmid, Neil Burch, Marc Lanctot, Matej Moravcik 외

Learning strategies for imperfect information games from samples of interaction is a challenging problem. A common method for this setting, Monte Carlo Counterfactual Regret Minimization (MCCFR), can have slow long-term …

counterfactualFormReinforcement Learning