paper-with-me

Papers

Provably Efficient Third-Person Imitation from Offline Observation

2020-02-27 · Aaron Zweig, Joan Bruna

Domain adaptation in imitation learning represents an essential step towards improving generalizability. However, even in the restricted setting of third-person imitation where transfer is between isomorphic Markov Decision Processes, there are no strong guarantees on the performance of transferred policies. We present problem-dependent, statistical learning guarantees for third-person imitation from observation in an offline setting, and a lower bound on performance in the online setting.

📄 PDF Abstract BibTeX arXiv:2002.12446

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationImitation Learning

Similar Papers 제목 키워드 기반

Online vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots

2024-09-12 · Ekaterina Svikhnushina, Pearl Pu

This paper explores the efficacy of online versus offline evaluation methods in assessing conversational chatbots, specifically comparing first-party direct interactions with third-party observational assessments. By ext…

BenchmarkingChatbot

Provably Efficient Causal Reinforcement Learning with Confounded Observational Data

2020-06-22 · NeurIPS 2021 12 · Lingxiao Wang, Zhuoran Yang, Zhaoran Wang

Empowered by expressive function approximators such as neural networks, deep reinforcement learning (DRL) achieves tremendous empirical successes. However, learning expressive function approximators requires collecting a…

Autonomous DrivingDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1

ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation

2025-09-23 · Jason Chen, I-Chun Arthur Liu, Gaurav Sukhatme, Daniel Seita arxiv

Training robust bimanual manipulation policies via imitation learning requires demonstration data with broad coverage over robot poses, contacts, and scene contexts. However, collecting diverse and precise real-world dem…

Data Augmentation

Third-Person Visual Imitation Learning via Decoupled Hierarchical Controller

2019-11-21 · NeurIPS 2019 12 · Pratyusha Sharma, Deepak Pathak, Abhinav Gupta

We study a generalized setup for learning from demonstration to build an agent that can manipulate novel objects in unseen scenarios by looking at only a single video of human demonstration from a third-person perspectiv…

Imitation Learning

Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning

2021-02-19 · Luofeng Liao, Zuyue Fu, Zhuoran Yang, Yixin Wang 외

In offline reinforcement learning (RL) an optimal policy is learned solely from a priori collected observational data. However, in observational data, actions are often confounded by unobserved variables. Instrumental va…

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1