paper-with-me

홈 › Papers

Policy-Controlled Generalized Share: A General Framework with a Transformer Instantiation for Strictly Online Switching-Oracle Tracking

2026-03-30 · Hongkai Hu arxiv

Static regret to a single expert is often the wrong target for strictly online prediction under non-stationarity, where the best expert may switch repeatedly over time. We study Policy-Controlled Generalized Share (PCGS), a general strictly online framework in which the generalized-share recursion is fixed while the post-loss update controls are allowed to vary adaptively. Its principal instantiation in this paper is PCGS-TF, which uses a causal Transformer as an update controller: after round t finishes and the loss vector is observed, the Transformer outputs the controls that map w_t to w_{t+1} without altering the already committed decision w_t. Under admissible post-loss update controls, we obtain a pathwise weighted regret guarantee for general time-varying learning rates, and a standard dynamic-regret guarantee against any expert path with at most S switches under the constant-learning-rate specialization. Empirically, on a controlled synthetic suite with exact dynamic-programming switching-oracle evaluation, PCGS-TF attains the lowest mean dynamic regret in all seven non-stationary families, with its advantage increasing for larger expert pools. On a reproduced household-electricity benchmark, PCGS-TF also achieves the lowest normalized dynamic regret for S = 5, 10, and 20.

📄 PDF Abstract BibTeX arXiv:2603.28198

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Preference Optimization by Estimating the Ratio of the Data Distribution

2025-05-26 · Yeongmin Kim, HeeSun Bae, Byeonghu Na, Il-Chul Moon

Direct preference optimization (DPO) is widely used as a simple and stable method for aligning large language models (LLMs) with human preferences. This paper investigates a generalized DPO loss that enables a policy mod…

Relaxed Actor-Critic with Convergence Guarantees for Continuous-Time Optimal Control of Nonlinear Systems

2019-09-11 · Jingliang Duan, Jie Li, Qiang Ge, Shengbo Eben Li 외

This paper presents the Relaxed Continuous-Time Actor-critic (RCTAC) algorithm, a method for finding the nearly optimal policy for nonlinear continuous-time (CT) systems with known dynamics and infinite horizon, such as …

Analysis of human steering behavior differences in human-in-control and autonomy-in-control driving

2024-09-30 · Rene Mai, Agung Julius, Sandipan Mishra

Steering models (such as the generalized two-point model) predict human steering behavior well when the human is in direct control of a vehicle. In vehicles under autonomous control, human control inputs are not used; ra…

State Estimation

Stability of Stochastic Approximations with `Controlled Markov' Noise and Temporal Difference Learning

2015-04-23 · Arunselvan Ramaswamy, Shalabh Bhatnagar

We are interested in understanding stability (almost sure boundedness) of stochastic approximation algorithms (SAs) driven by a `controlled Markov' process. Analyzing this class of algorithms is important, since many rei…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Reinforcement Learning of Control Policy for Linear Temporal Logic Specifications Using Limit-Deterministic Generalized Büchi Automata

2020-01-14 · Ryohei Oura, Ami Sakakibara, Toshimitsu Ushio

This letter proposes a novel reinforcement learning method for the synthesis of a control policy satisfying a control specification described by a linear temporal logic formula. We assume that the controlled system is mo…

Reinforcement Learning