paper-with-me

홈 › Papers

Can CDT rationalise the ex ante optimal policy via modified anthropics?

2024-11-07 · Emery Cooper, Caspar Oesterheld, Vincent Conitzer

In Newcomb's problem, causal decision theory (CDT) recommends two-boxing and thus comes apart from evidential decision theory (EDT) and ex ante policy optimisation (which prescribe one-boxing). However, in Newcomb's problem, you should perhaps believe that with some probability you are in a simulation run by the predictor to determine whether to put a million dollars into the opaque box. If so, then causal decision theory might recommend one-boxing in order to cause the predictor to fill the opaque box. In this paper, we study generalisations of this approach. That is, we consider general Newcomblike problems and try to form reasonable self-locating beliefs under which CDT's recommendations align with an EDT-like notion of ex ante policy optimisation. We consider approaches in which we model the world as running simulations of the agent, and an approach not based on such models (which we call 'Generalised Generalised Thirding', or GGT). For each approach, we characterise the resulting CDT policies, and prove that under certain conditions, these include the ex ante optimal policies.

📄 PDF Abstract BibTeX arXiv:2411.04462

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation

2024-11-15 · Yihong Guo, YiXuan Wang, Yuanyuan Shi, Pan Xu 외

Training a policy in a source domain for deployment in the target domain under a dynamics shift can be challenging, often resulting in performance degradation. Previous work tackles this challenge by training on the sour…

Domain AdaptationImitation Learning

Tight Performance Bounds for Approximate Modified Policy Iteration with Non-Stationary Policies

2013-04-20 · Boris Lesner, Bruno Scherrer

We consider approximate dynamic programming for the infinite-horizon stationary $\gamma$-discounted optimal control problem formalized by Markov Decision Processes. While in the exact case it is known that there always e…

Beyond the Policy Gradient Theorem for Efficient Policy Updates in Actor-Critic Algorithms

2022-02-15 · Romain Laroche, Remi Tachet

In Reinforcement Learning, the optimal action at a given state is dependent on policy decisions at subsequent states. As a consequence, the learning targets evolve with time and the policy optimization process must be ef…

Constrained Best Arm Identification in Grouped Bandits

2024-12-11 · Sahil Dharod, Malyala Preethi Sravani, Sakshi Heda, Sharayu Moharir

We study a grouped bandit setting where each arm comprises multiple independent sub-arms referred to as attributes. Each attribute of each arm has an independent stochastic reward. We impose the constraint that for an ar…

Attribute

Models and algorithms for skip-free Markov decision processes on trees

2013-09-17 · E. J. Collins

We introduce a class of models for multidimensional control problems which we call skip-free Markov decision processes on trees. We describe and analyse an algorithm applicable to Markov decision processes of this type t…