paper-with-me

Papers

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

2026-06-24 · Guozheng Ma, Lu Li, Zilin Wang, Pierre-Luc Bacon, Dacheng Tao arxiv

Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offline-to-online RL, this paradigm now spans foundation model post-training and embodied intelligence, with prior types expanding from offline datasets and pre-trained policies to increasingly diverse knowledge sources such as multimodal foundation models and generative world models. Offline priors have become central to how deep RL is developed and deployed. However, this reliance introduces a challenge that the prevailing benchmark-driven paradigm cannot resolve: because prior validity varies across deployments and shifts during training, no single approach to managing it is universally optimal, and benchmark rankings offer limited guidance for real-world deployments. Rather than pursuing universal solutions, we argue that the field should shift to diagnosis-driven tension management, in which deployment-specific evidence guides how the learner relates to its priors throughout training, enabling both flexible and adaptive deployment. We support this position with a framework characterizing how priors reshape online optimization through three functional roles, controlled experiments demonstrating help-or-hurt reversals, cross-domain evidence from foundation model post-training to embodied intelligence, and engagement with five substantive counterarguments.

📄 PDF Abstract BibTeX arXiv:2606.25527

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning

2023-06-14 · Nikhil Vyas, Depen Morwani, Rosie Zhao, Gal Kaplun 외

The success of SGD in deep learning has been ascribed by prior works to the implicit bias induced by finite batch sizes ("SGD noise"). While prior works focused on offline learning (i.e., multiple-epoch training), we stu…

Fault diagnosis for open-circuit faults in NPC inverter based on knowledge-driven and data-driven approaches

2022-10-31 · Lei Kou, Chuang Liu, Guo-wei Cai, Jia-ning Zhou 외

In this study, the open-circuit faults diagnosis and location issue of the neutral-point-clamped (NPC) inverters are analysed. A novel fault diagnosis approach based on knowledge driven and data driven was presented for …

Fault Diagnosis

Data-driven design of fault diagnosis for three-phase PWM rectifier using random forests technique with transient synthetic features

2022-11-02 · Lei Kou, Chuang Liu, Guo-wei Cai, Jia-ning Zhou 외

A three-phase pulse-width modulation (PWM) rectifier can usually maintain operation when open-circuit faults occur in insulated-gate bipolar transistors (IGBTs), which will lead the system to be unstable and unsafe. Aimi…

Fault Diagnosis

Task-oriented Dialogue System for Automatic Diagnosis

2018-07-01 · ACL 2018 7 · Zhongyu Wei, Qianlong Liu, Baolin Peng, Huaixiao Tou 외

In this paper, we make a move to build a dialogue system for automatic diagnosis. We first build a dataset collected from an online medical forum by extracting symptoms from both patients{'} self-reports and conversation…

DeepPolar: Inventing Nonlinear Large-Kernel Polar Codes via Deep Learning

2024-02-14 · S Ashwin Hebbar, Sravan Kumar Ankireddy, Hyeji Kim, Sewoong Oh 외

Progress in designing channel codes has been driven by human ingenuity and, fittingly, has been sporadic. Polar codes, developed on the foundation of Arikan's polarization kernel, represent the latest breakthrough in cod…

Deep LearningIngenuity