paper-with-me

Papers

Double Check Your State Before Trusting It: Confidence-Aware Bidirectional Offline Model-Based Imagination

2022-06-16 · Jiafei Lyu, Xiu Li, Zongqing Lu

The learned policy of model-free offline reinforcement learning (RL) methods is often constrained to stay within the support of datasets to avoid possible dangerous out-of-distribution actions or states, making it challenging to handle out-of-support region. Model-based RL methods offer a richer dataset and benefit generalization by generating imaginary trajectories with either trained forward or reverse dynamics model. However, the imagined transitions may be inaccurate, thus downgrading the performance of the underlying offline RL method. In this paper, we propose to augment the offline dataset by using trained bidirectional dynamics models and rollout policies with double check. We introduce conservatism by trusting samples that the forward model and backward model agree on. Our method, confidence-aware bidirectional offline model-based imagination, generates reliable samples and can be combined with any model-free offline RL method. Experimental results on the D4RL benchmarks demonstrate that our method significantly boosts the performance of existing model-free offline RL algorithms and achieves competitive or better scores against baseline methods.

📄 PDF Abstract BibTeX arXiv:2206.07989

Code (1)

dmksjfl/CABI 공식 구현 pytorch

Tasks

D4RLOffline RLReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

LeanCSP: A Framework for Certifying Constraint Reformulation and Solving in Lean

2026-07-30 · Pablo Manrique, Stefan Szeider arxiv

Constraint programming is a core technology for solving complex combinatorial problems in scheduling, planning, configuration, and verification. Trusting its results therefore demands guarantees at two levels: that refor…

"Image, Tell me your story!" Predicting the original meta-context of visual misinformation

2024-08-19 · Jonathan Tonglet, Marie-Francine Moens, Iryna Gurevych

To assist human fact-checkers, researchers have developed automated approaches for visual misinformation detection. These methods assign veracity scores by identifying inconsistencies between the image and its caption, o…

Fact CheckingMisinformation

Trusting RoBERTa over BERT: Insights from CheckListing the Natural Language Inference Task

2021-07-15 · Ishan Tarunesh, Somak Aditya, Monojit Choudhury

The recent state-of-the-art natural language understanding (NLU) systems often behave unpredictably, failing on simpler reasoning examples. Despite this, there has been limited focus on quantifying progress towards syste…

Natural Language InferenceNatural Language Understanding

Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning

2025-06-26 · Xin Xu, Tianhao Chen, Fan Zhang, Wanlong Liu 외

While slow-thinking large language models (LLMs) exhibit reflection-like reasoning, commonly referred to as the "aha moment:, their ability to generate informative critiques and refine prior solutions remains limited. In…

CheckYourMeal!: diet management with NLG

2018-11-01 · WS 2018 11 · Luca Anselma, Simone Donetti, Aless Mazzei, ro 외
ManagementText Generation