Papers Q-Learning
“Q-Learning” 태그가 달린 논문 1,918편 · 필터 해제
Evaluating Reinforcement Learning Algorithms for Navigation in Simulated Robotic Quadrupeds: A Comparative Study Inspired by Guide Dog Behaviour
Robots are increasingly integrated across industries, particularly in healthcare. However, many valuable applications for quadrupedal robots remain overlooked. This research explores the effectiveness of three reinforcem…
Autonomous NavigationQ-LearningPersonalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing
We introduce ExRec, a general framework for personalized exercise recommendation with semantically-grounded knowledge tracing. Our method builds on the observation that existing exercise recommendation approaches simulat…
Knowledge TracingMathQ-LearningReinforcement Learning (RL)A Data-Ensemble-Based Approach for Sample-Efficient LQ Control of Linear Time-Varying Systems
This paper presents a sample-efficient, data-driven control framework for finite-horizon linear quadratic (LQ) control of linear time-varying (LTV) systems. In contrast to the time-invariant case, the time-varying LQ pro…
Q-LearningADDQ: Adaptive Distributional Double Q-Learning
Bias problems in the estimation of $Q$-values are a well-known obstacle that slows down convergence of $Q$-learning and actor-critic methods. One of the reasons of the success of modern RL algorithms is partially a direc…
Distributional Reinforcement LearningMuJoCoQ-LearningReinforcement Learning-Based Policy Optimisation For Heterogeneous Radio Access
Flexible and efficient wireless resource sharing across heterogeneous services is a key objective for future wireless networks. In this context, we investigate the performance of a system where latency-constrained intern…
Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)ReinDSplit: Reinforced Dynamic Split Learning for Pest Recognition in Precision Agriculture
To empower precision agriculture through distributed machine learning (DML), split learning (SL) has emerged as a promising paradigm, partitioning deep neural networks (DNNs) between edge devices and servers to reduce co…
Q-LearningReinforcement Learning (RL)Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning
Offline reinforcement learning promises policy improvement from logged interaction data alone, yet state-of-the-art algorithms remain vulnerable to value over-estimation and to violations of domain knowledge such as mono…
Q-Learning"What are my options?": Explaining RL Agents with Diverse Near-Optimal Alternatives (Extended)
In this work, we provide an extended discussion of a new approach to explainable Reinforcement Learning called Diverse Near-Optimal Alternatives (DNA), first proposed at L4DC 2025. DNA seeks a set of reasonable "options"…
DiversityQ-LearningStochastic OptimizationTrajectory PlanningQ-learning-based Hierarchical Cooperative Local Search for Steelmaking-continuous Casting Scheduling Problem
The steelmaking continuous casting scheduling problem (SCCSP) is a critical and complex challenge in modern steel production, requiring the coordinated assignment and sequencing of steel charges across multiple productio…
Q-LearningSchedulingRegret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning
Motivated by real-world settings where data collection and policy deployment -- whether for a single agent or across multiple agents -- are costly, we study the problem of on-policy single-agent reinforcement learning (R…
Q-LearningReinforcement Learning (RL)Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning With Iterated Q-Learning
In value-based reinforcement learning, removing the target network is tempting as the boostrapped target would be built from up-to-date estimates, and the spared memory occupied by the target network could be reallocated…
Q-LearningImproving Performance of Spike-based Deep Q-Learning using Ternary Neurons
We propose a new ternary spiking neuron model to improve the representation capacity of binary spiking neurons in deep Q-learning. Although a ternary neuron model has recently been introduced to overcome the limited repr…
Atari GamesDecision MakingQ-LearningReinforcement Learning for Hanabi
Hanabi has become a popular game for research when it comes to reinforcement learning (RL) as it is one of the few cooperative card games where you have incomplete knowledge of the entire environment, thus presenting a c…
Card GamesDeep Reinforcement LearningQ-Learningreinforcement-learning+2Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model
In this paper we analyze the sample complexities of learning the optimal state-action value function $Q^*$ and an optimal policy $\pi^*$ in a discounted Markov decision process (MDP) where the agent has recursive entropi…
Q-LearningOn Global Convergence Rates for Federated Policy Gradient under Heterogeneous Environment
Ensuring convergence of policy gradient methods in federated reinforcement learning (FRL) under environment heterogeneity remains a major challenge. In this work, we first establish that heterogeneity, perhaps counter-in…
Federated LearningPolicy Gradient MethodsQ-LearningBOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL
Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in the context of single-objective BO, lear…
Bayesian OptimizationHyperparameter OptimizationLanguage ModelingLanguage Modelling+1Learning to Charge More: A Theoretical Study of Collusion by Q-Learning Agents
There is growing experimental evidence that $Q$-learning agents may learn to charge supracompetitive prices. We provide the first theoretical explanation for this behavior in infinite repeated games. Firms update their p…
Q-LearningA General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging
Polyak-Ruppert averaging is a widely used technique to achieve the optimal asymptotic variance of stochastic approximation (SA) algorithms, yet its high-probability performance guarantees remain underexplored in general …
Q-LearningInverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
We study the problem of offline imitation learning in Markov decision processes (MDPs), where the goal is to learn a well-performing policy given a dataset of state-action pairs generated by an expert policy. Complementi…
Imitation LearningQ-LearningDistributionally Robust Deep Q-Learning
We propose a novel distributionally robust $Q$-learning algorithm for the non-tabular case accounting for continuous state spaces where the state transition of the underlying Markov decision process is subject to model u…
Q-Learning