paper-with-me

홈 › Papers

SOAC: The Soft Option Actor-Critic Architecture

2020-06-25 · Chenghao Li, Xiaoteng Ma, Chongjie Zhang, Jun Yang, Li Xia, Qianchuan Zhao

The option framework has shown great promise by automatically extracting temporally-extended sub-tasks from a long-horizon task. Methods have been proposed for concurrently learning low-level intra-option policies and high-level option selection policy. However, existing methods typically suffer from two major challenges: ineffective exploration and unstable updates. In this paper, we present a novel and stable off-policy approach that builds on the maximum entropy model to address these challenges. Our approach introduces an information-theoretical intrinsic reward for encouraging the identification of diverse and effective options. Meanwhile, we utilize a probability inference model to simplify the optimization problem as fitting optimal trajectories. Experimental results demonstrate that our approach significantly outperforms prior on-policy and off-policy methods in a range of Mujoco benchmark tasks while still providing benefits for transfer learning. In these tasks, our approach learns a diverse set of options, each of whose state-action space has strong coherence.

📄 PDF Abstract BibTeX arXiv:2006.14363

Code (0)

등록된 구현이 없습니다.

Tasks

MuJoCoTransfer Learning

Similar Papers 제목 키워드 기반

Soft Options Critic

2019-05-23 · Elita Lobo, Scott Jordan

The option-critic architecture (Bacon, Harb, and Precup 2017) and several variants have successfully demonstrated the use of the options framework proposed by Sutton et al (Sutton, Precup, and Singh1999) to scale learnin…

Predicting Individual Responses to Vasoactive Medications in Children with Septic Shock

2019-01-15 · Nicole Fronda, Jessica Asencio, Cameron Carlin, David Ledbetter 외

Objective: Predict individual septic children's personalized physiologic responses to vasoactive titrations by training a Recurrent Neural Network (RNN) using EMR data. Materials and Methods: This study retrospectively…

Holdout SetregressionTime Series Analysis

DAC: The Double Actor-Critic Architecture for Learning Options

2019-04-29 · NeurIPS 2019 12 · Shangtong Zhang, Shimon Whiteson

We reformulate the option framework as two parallel augmented MDPs. Under this novel formulation, all policy optimization algorithms can be used off the shelf to learn intra-option policies, option termination conditions…

Transfer Learning

A Novel Multi-Task Teacher-Student Architecture with Self-Supervised Pretraining for 48-Hour Vasoactive-Inotropic Trend Analysis in Sepsis Mortality Prediction

2025-02-24 · Houji Jin, Negin Ashrafi, Kamiar Alaei, Elham Pishgar 외

Sepsis is a major cause of ICU mortality, where early recognition and effective interventions are essential for improving patient outcomes. However, the vasoactive-inotropic score (VIS) varies dynamically with a patient'…

ICU MortalityMortality Prediction

From Expectation to Habit: Why Do Software Practitioners Adopt Fairness Toolkits?

2024-12-18 · Gianmario Voria, Stefano Lambiase, Maria Concetta Schiavone, Gemma Catolino 외

As the adoption of machine learning (ML) systems continues to grow across industries, concerns about fairness and bias in these systems have taken center stage. Fairness toolkits, designed to mitigate bias in ML models, …

Fairness