paper-with-me

Papers

Learning Task Decomposition with Order-Memory Policy Network

2021-01-01 · ICLR 2021 1 · Yuchen Lu, Yikang Shen, Siyuan Zhou, Aaron Courville, Joshua B. Tenenbaum, Chuang Gan

Many complex real-world tasks are composed of several levels of sub-tasks. Humans leverage these hierarchical structures to accelerate the learning process and achieve better generalization. To simulate this process, we introduce Ordered Memory Policy Network (OMPN) to discover task decomposition by imitation learning from demonstration. OMPN has an explicit inductive bias to model a hierarchy of sub-tasks. Experiments on Craft world and Dial demonstrate that our model can more accurately recover the task boundaries with behavior cloning under both unsupervised and weakly supervised setting than previous methods. OMPN can also be directly applied to partially observable environments and still achieve high performance. Our visualization further confirms the intuition that OMPN can learn to expand the memory at higher levels when one subtask is close to completion.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation LearningInductive Bias

Similar Papers 제목 키워드 기반

Learning Task Decomposition with Ordered Memory Policy Network

2021-03-19 · Yuchen Lu, Yikang Shen, Siyuan Zhou, Aaron Courville 외

Many complex real-world tasks are composed of several levels of sub-tasks. Humans leverage these hierarchical structures to accelerate the learning process and achieve better generalization. In this work, we study the in…

Inductive Bias

Activation Map Compression through Tensor Decomposition for Deep Learning

2024-11-10 · Le-Trung Nguyen, Aël Quélennec, Enzo Tartaglione, Samuel Tardieu 외

Internet of Things and Deep Learning are synergetically and exponentially growing industrial fields with a massive call for their unification into a common framework called Edge AI. While on-device inference is a well-ex…

Deep LearningTensor Decomposition

Reinforcement Learning of POMDPs using Spectral Methods

2016-02-25 · Kamyar Azizzadenesheli, Alessandro Lazaric, Animashree Anandkumar

We propose a new reinforcement learning algorithm for partially observable Markov decision processes (POMDP) based on spectral decomposition methods. While spectral methods have been previously employed for consistent le…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Experimental results : Reinforcement Learning of POMDPs using Spectral Methods

2017-05-07 · Kamyar Azizzadenesheli, Alessandro Lazaric, Animashree Anandkumar

We propose a new reinforcement learning algorithm for partially observable Markov decision processes (POMDP) based on spectral decomposition methods. While spectral methods have been previously employed for consistent le…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Shape of Memory: a Geometric Analysis of Machine Unlearning in Second-Order Optimizers

2026-04-24 · Kennon Stewart arxiv

We argue that current definitions of machine unlearning are underspecified for second-order optimizers. We compare first-order and second-order learners for their ability to handle the data deletion task with varying deg…