paper-with-me

홈 › Papers

A Critique of Strictly Batch Imitation Learning

2021-10-05 · Gokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Zhiwei Steven Wu

Recent work by Jarrett et al. attempts to frame the problem of offline imitation learning (IL) as one of learning a joint energy-based model, with the hope of out-performing standard behavioral cloning. We suggest that notational issues obscure how the psuedo-state visitation distribution the authors propose to optimize might be disconnected from the policy's $\textit{true}$ state visitation distribution. We further construct natural examples where the parameter coupling advocated by Jarrett et al. leads to inconsistent estimates of the expert's policy, unlike behavioral cloning.

📄 PDF Abstract BibTeX arXiv:2110.02063

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation Learning

Similar Papers 제목 키워드 기반

Strictly Batch Imitation Learning by Energy-based Distribution Matching

2020-06-25 · NeurIPS 2020 12 · Daniel Jarrett, Ioana Bica, Mihaela van der Schaar

Consider learning a policy purely on the basis of demonstrated behavior -- that is, with no access to reinforcement signals, no knowledge of transition dynamics, and no further interaction with the environment. This *str…

Imitation LearningOff-policy evaluation

A Critique of a Critique of Word Similarity Datasets: Sanity Check or Unnecessary Confusion?

2017-07-12 · Minh Le

Critical evaluation of word similarity datasets is very important for computational lexical semantics. This short report concerns the sanity check proposed in Batchkarov et al. (2016) to evaluate several popular datasets…

Word Similarity

Incremental Clustering: The Case for Extra Clusters

2014-06-24 · NeurIPS 2014 12 · Margareta Ackerman, Sanjoy Dasgupta

The explosion in the amount of data available for analysis often necessitates a transition from batch to incremental clustering methods, which process one element at a time and typically store only a small subset of the …

Clustering

CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

2025-02-23 · Alexander Zhang, Marcus Dong, Jiaheng Liu, Wei zhang 외

The critique capacity of Large Language Models (LLMs) is essential for reasoning abilities, which can provide necessary suggestions (e.g., detailed analysis and constructive feedback). Therefore, how to evaluate the crit…

Code GenerationHumanEvalmbpp

Shepherd: A Critic for Language Model Generation

2023-08-08 · Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan, Sean O'Brien 외

As large language models improve, there is increasing interest in techniques that leverage these models' capabilities to refine their own outputs. In this work, we introduce Shepherd, a language model specifically tuned …

Language ModelingLanguage Modellingmodel