paper-with-me

Papers

Flexible Option Learning

2021-12-06 · NeurIPS 2021 12 · Martin Klissarov, Doina Precup

Temporal abstraction in reinforcement learning (RL), offers the promise of improving generalization and knowledge transfer in complex environments, by propagating information more efficiently over time. Although option learning was initially formulated in a way that allows updating many options simultaneously, using off-policy, intra-option learning (Sutton, Precup & Singh, 1999), many of the recent hierarchical reinforcement learning approaches only update a single option at a time: the option currently executing. We revisit and extend intra-option learning in the context of deep reinforcement learning, in order to enable updating all options consistent with current primitive action choices, without introducing any additional estimates. Our method can therefore be naturally adopted in most hierarchical RL frameworks. When we combine our approach with the option-critic algorithm for option discovery, we obtain significant improvements in performance and data-efficiency across a wide variety of domains.

📄 PDF Abstract BibTeX arXiv:2112.03097

Code (1)

mklissa/moc 공식 구현 tf

Tasks

Deep Reinforcement LearningHierarchical Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Transfer Learning

Similar Papers 제목 키워드 기반

European Option Pricing in Regime Switching Framework via Physics-Informed Residual Learning

2024-10-14 · Naman Krishna Pande, Puneet Pasricha, Arun Kumar, Arvind Kumar Gupta

In this article, we employ physics-informed residual learning (PIRL) and propose a pricing method for European options under a regime-switching framework, where closed-form solutions are not available. We demonstrate tha…

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering

2026-04-18 · Xiyin Zeng, Yi Lu, Hao Wang arxiv

Visual Question Answering (VQA) requires models to identify the correct answer options based on both visual and textual evidence. Recent Mixture-of-Experts (MoE) methods improve option reasoning by grouping similar conce…

Visual Question AnsweringContrastive Learning

The least squares method for option pricing revisited

2015-11-16

It is shown that the the popular least squares method of option pricing converges even under very general assumptions. This substantially increases the freedom of creating different implementations of the method, with va…

regression

Spiking Nonlinear Opinion Dynamics (S-NOD) for Agile Decision-Making

2024-09-18 · Charlotte Cathcart, Ian Xul Belaustegui, Alessio Franci, Naomi Ehrich Leonard

We present, analyze, and illustrate a first-of-its-kind model of two-dimensional excitable (spiking) dynamics for decision-making over two options. The model, Spiking Nonlinear Opinion Dynamics (S-NOD), provides superior…

Decision MakingRobot Navigation

Exchangeable Gaussian Processes for Staggered-Adoption Policy Evaluation

2026-02-24 · Hayk Gevorgyan, Konstantinos Kalogeropoulos, Angelos Alexopoulos arxiv

We study the use of exchangeable multi-task Gaussian processes (GPs) for causal inference in panel data, applying the framework to two settings: one with a single treated unit subject to a once-and-for-all treatment and …

Gaussian ProcessesCausal Inference