paper-with-me

홈 › Papers

SMILe: Scalable Meta Inverse Reinforcement Learning through Context-Conditional Policies

2019-12-01 · NeurIPS 2019 12 · Seyed Kamyar Seyed Ghasemipour, Shixiang (Shane) Gu, Richard Zemel

Imitation Learning (IL) has been successfully applied to complex sequential decision-making problems where standard Reinforcement Learning (RL) algorithms fail. A number of recent methods extend IL to few-shot learning scenarios, where a meta-trained policy learns to quickly master new tasks using limited demonstrations. However, although Inverse Reinforcement Learning (IRL) often outperforms Behavioral Cloning (BC) in terms of imitation quality, most of these approaches build on BC due to its simple optimization objective. In this work, we propose SMILe, a scalable framework for Meta Inverse Reinforcement Learning (Meta-IRL) based on maximum entropy IRL, which can learn high-quality policies from few demonstrations. We examine the efficacy of our method on a variety of high-dimensional simulated continuous control tasks and observe that SMILe significantly outperforms Meta-BC. Furthermore, we observe that SMILe performs comparably or outperforms Meta-DAgger, while being applicable in the state-only setting and not requiring online experts. To our knowledge, our approach is the first efficient method for Meta-IRL that scales to the function approximator setting. For datasets and reproducing results please refer to https://github.com/KamyarGh/rl_swiss/blob/master/reproducing/smile_paper.md .

📄 PDF Abstract BibTeX

Code (1)

KamyarGh/rl_swiss 공식 구현 pytorch

Tasks

continuous-controlContinuous ControlDecision MakingFew-Shot LearningImitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Sequential Decision Making

Similar Papers 제목 키워드 기반

Fast and scalable retrosynthetic planning with a transformer neural network and speculative beam search

2025-08-02 · Mikhail Andronov, Natalia Andronova, Michael Wand, Jürgen Schmidhuber 외 arxiv

AI-based computer-aided synthesis planning (CASP) systems are in demand as components of AI-driven drug discovery workflows. However, the high latency of such CASP systems limits their utility for high-throughput synthes…

Single-step retrosynthesisDrug Discovery

Meta-Cognition. An Inverse-Inverse Reinforcement Learning Approach for Cognitive Radars

2022-05-03 · Kunal Pattanayak, Vikram Krishnamurthy, Christopher Berry

This paper considers meta-cognitive radars in an adversarial setting. A cognitive radar optimally adapts its waveform (response) in response to maneuvers (probes) of a possibly adversarial moving target. A meta-cognitive…

reinforcement-learningReinforcement Learning (RL)

SP2RINT: Spatially-Decoupled Physics-Inspired Progressive Inverse Optimization for Scalable, PDE-Constrained Meta-Optical Neural Network Training

2025-05-23 · Pingchuan Ma, Ziang Yin, Qi Jing, Zhengqi Gao 외

DONNs leverage light propagation for efficient analog AI and signal processing. Advances in nanophotonic fabrication and metasurface-based wavefront engineering have opened new pathways to realize high-capacity DONNs acr…

SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition

2024-09-16 · Ming-Hao Hsu, Hung-Yi Lee

Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource languages but face significant challenges when applied to low-resource languages due to limited training data and insufficien…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationIn-Context Learning+4

Can Large Language Models Understand Molecules?

2024-01-05 · Shaghayegh Sadeghi, Alan Bui, Ali Forooghi, Jianguo Lu 외

Purpose: Large Language Models (LLMs) like GPT (Generative Pre-trained Transformer) from OpenAI and LLaMA (Large Language Model Meta AI) from Meta AI are increasingly recognized for their potential in the field of chemin…

Drug DiscoveryLanguage ModellingLarge Language ModelMolecular Property Prediction+3