paper-with-me

홈 › Papers

Optimism is All You Need: Model-Based Imitation Learning From Observation Alone

2021-03-09 · ICLR Workshop SSL-RL 2021 5 · Rahul Kidambi, Jonathan Daniel Chang, Wen Sun

This paper studies Imitation Learning from Observations alone (ILFO) where the learner is presented with expert demonstrations that only consist of states encountered by an expert (without access to actions taken by the expert). This paper presents a provably efficient model-based framework MobILE to solve the ILFO problem. MobILE uses self-supervision towards (a) training a dynamics model and (b) designing an intrinsic reward signal for exploration. Using these ideas, MobILE carefully trades off exploration against imitation by integrating the idea of optimism in the face of uncertainty into the distribution matching imitation learning (IL) framework. We provide a unified analysis for MobILE, and demonstrate that MobILE enjoys strong performance guarantees for classes of MDP dynamics that satisfy certain well studied notions of complexity. We also show that the ILFO problem is strictly harder than the standard IL problem by reducing ILFO to a multi-armed bandit problem indicating that strategic exploration is necessary for solving ILFO efficiently. We complement these theoretical results with experimental simulations on benchmark OpenAI Gym tasks that indicate the efficacy of MobILE.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

AllImitation LearningOpenAI Gym

Similar Papers 제목 키워드 기반

MobILE: Model-Based Imitation Learning From Observation Alone

2021-02-22 · NeurIPS 2021 12 · Rahul Kidambi, Jonathan Chang, Wen Sun

This paper studies Imitation Learning from Observations alone (ILFO) where the learner is presented with expert demonstrations that consist only of states visited by an expert (without access to actions taken by the expe…

Imitation LearningmodelOpenAI Gym

Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification

2025-11-28 · Saeed AlMarri, Mathieu Ravaut, Kristof Juhasz, Gautier Marti 외 arxiv

Large Language Models (LLMs) have attracted significant attention for classification tasks, offering a flexible alternative to trusted classical machine learning models like LightGBM through zero-shot prompting. However,…

Provably Efficient Imitation Learning from Observation Alone

2019-05-27 · Wen Sun, Anirudh Vemula, Byron Boots, J. Andrew Bagnell

We study Imitation Learning (IL) from Observations alone (ILFO) in large-scale MDPs. While most IL algorithms rely on an expert to directly provide actions to the learner, in this setting the expert only supplies sequenc…

Imitation LearningOpenAI GymReinforcement Learning

Supervised Optimism Correction: Be Confident When LLMs Are Sure

2025-04-10 · Junjie Zhang, Rushuai Yang, Shunyu Liu, Ting-En Lin 외

In this work, we establish a novel theoretical connection between supervised fine-tuning and offline reinforcement learning under the token-level Markov decision process, revealing that large language models indeed learn…

GSM8KMathMathematical Reasoning

Exploring Optimism and Pessimism in Twitter Using Deep Learning

2018-10-01 · EMNLP 2018 10 · Cornelia Caragea, Liviu P. Dinu, Bogdan Dumitru

Identifying optimistic and pessimistic viewpoints and users from Twitter is useful for providing better social support to those who need such support, and for minimizing the negative influence among users and maximizing …

Deep Learning