paper-with-me

Papers Q-Learning

“Q-Learning” 태그가 달린 논문 1,918편 · 필터 해제

Evaluating Reinforcement Learning Algorithms for Navigation in Simulated Robotic Quadrupeds: A Comparative Study Inspired by Guide Dog Behaviour

2025-07-17 · Emma M. A. Harrison

Robots are increasingly integrated across industries, particularly in healthcare. However, many valuable applications for quadrupedal robots remain overlooked. This research explores the effectiveness of three reinforcem…

Autonomous NavigationQ-Learning

Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing

2025-07-15 · Yilmazcan Ozyurt, Tunaberk Almaci, Stefan Feuerriegel, Mrinmaya Sachan

We introduce ExRec, a general framework for personalized exercise recommendation with semantically-grounded knowledge tracing. Our method builds on the observation that existing exercise recommendation approaches simulat…

Knowledge TracingMathQ-LearningReinforcement Learning (RL)

A Data-Ensemble-Based Approach for Sample-Efficient LQ Control of Linear Time-Varying Systems

2025-06-30 · Sahel Vahedi Noori, Maryam Babazadeh

This paper presents a sample-efficient, data-driven control framework for finite-horizon linear quadratic (LQ) control of linear time-varying (LTV) systems. In contrast to the time-invariant case, the time-varying LQ pro…

Q-Learning

ADDQ: Adaptive Distributional Double Q-Learning

2025-06-24 · Leif Döring, Benedikt Wille, Maximilian Birr, Mihail Bîrsan 외

Bias problems in the estimation of $Q$-values are a well-known obstacle that slows down convergence of $Q$-learning and actor-critic methods. One of the reasons of the success of modern RL algorithms is partially a direc…

Distributional Reinforcement LearningMuJoCoQ-Learning

Reinforcement Learning-Based Policy Optimisation For Heterogeneous Radio Access

2025-06-18 · Anup Mishra, Čedomir Stefanović, Xiuqiang Xu, Petar Popovski 외

Flexible and efficient wireless resource sharing across heterogeneous services is a key objective for future wireless networks. In this context, we investigate the performance of a system where latency-constrained intern…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

ReinDSplit: Reinforced Dynamic Split Learning for Pest Recognition in Precision Agriculture

2025-06-16 · Vishesh Kumar Tanwar, Soumik Sarkar, Asheesh K. Singh, Sajal K. Das

To empower precision agriculture through distributed machine learning (DML), split learning (SL) has emerged as a promising paradigm, partitioning deep neural networks (DNNs) between edge devices and servers to reduce co…

Q-LearningReinforcement Learning (RL)

Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning

2025-06-16 · Ali Baheri

Offline reinforcement learning promises policy improvement from logged interaction data alone, yet state-of-the-art algorithms remain vulnerable to value over-estimation and to violations of domain knowledge such as mono…

Q-Learning

"What are my options?": Explaining RL Agents with Diverse Near-Optimal Alternatives (Extended)

2025-06-11 · Noel Brindise, Vijeth Hebbar, Riya Shah, Cedric Langbort

In this work, we provide an extended discussion of a new approach to explainable Reinforcement Learning called Diverse Near-Optimal Alternatives (DNA), first proposed at L4DC 2025. DNA seeks a set of reasonable "options"…

DiversityQ-LearningStochastic OptimizationTrajectory Planning

Q-learning-based Hierarchical Cooperative Local Search for Steelmaking-continuous Casting Scheduling Problem

2025-06-10 · Yang Lv, Rong Hu, Bin Qian, Jian-Bo Yang

The steelmaking continuous casting scheduling problem (SCCSP) is a critical and complex challenge in modern steel production, requiring the coordinated assignment and sequencing of steel charges across multiple productio…

Q-LearningScheduling

Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning

2025-06-05 · Haochen Zhang, Zhong Zheng, Lingzhou Xue

Motivated by real-world settings where data collection and policy deployment -- whether for a single agent or across multiple agents -- are costly, we study the problem of on-policy single-agent reinforcement learning (R…

Q-LearningReinforcement Learning (RL)

Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning With Iterated Q-Learning

2025-06-04 · Théo Vincent, Yogesh Tripathi, Tim Faust, Yaniv Oren 외

In value-based reinforcement learning, removing the target network is tempting as the boostrapped target would be built from up-to-date estimates, and the spared memory occupied by the target network could be reallocated…

Q-Learning

Improving Performance of Spike-based Deep Q-Learning using Ternary Neurons

2025-06-03 · Aref Ghoreishee, Abhishek Mishra, John Walsh, Anup Das 외

We propose a new ternary spiking neuron model to improve the representation capacity of binary spiking neurons in deep Q-learning. Although a ternary neuron model has recently been introduced to overcome the limited repr…

Atari GamesDecision MakingQ-Learning

Reinforcement Learning for Hanabi

2025-05-31 · Nina Cohen, Kordel K. France

Hanabi has become a popular game for research when it comes to reinforcement learning (RL) as it is one of the few cooperative card games where you have incomplete knowledge of the entire environment, thus presenting a c…

Card GamesDeep Reinforcement LearningQ-Learningreinforcement-learning+2

Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model

2025-05-30 · Oliver Mortensen, Mohammad Sadegh Talebi

In this paper we analyze the sample complexities of learning the optimal state-action value function $Q^*$ and an optimal policy $\pi^*$ in a discounted Markov decision process (MDP) where the agent has recursive entropi…

Q-Learning

On Global Convergence Rates for Federated Policy Gradient under Heterogeneous Environment

2025-05-29 · Safwan Labbi, Paul Mangold, Daniil Tiapkin, Eric Moulines

Ensuring convergence of policy gradient methods in federated reinforcement learning (FRL) under environment heterogeneity remains a major challenge. In this work, we first establish that heterogeneity, perhaps counter-in…

Federated LearningPolicy Gradient MethodsQ-Learning

BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL

2025-05-28 · Yu-Heng Hung, Kai-Jie Lin, Yu-Heng Lin, Chien-YiWang 외

Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in the context of single-objective BO, lear…

Bayesian OptimizationHyperparameter OptimizationLanguage ModelingLanguage Modelling+1

Learning to Charge More: A Theoretical Study of Collusion by Q-Learning Agents

2025-05-28 · Cristian Chica, Yinglong Guo, Gilad Lerman

There is growing experimental evidence that $Q$-learning agents may learn to charge supracompetitive prices. We provide the first theoretical explanation for this behavior in infinite repeated games. Firms update their p…

Q-Learning

A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging

2025-05-27 · Sajad Khodadadian, Martin Zubeldia

Polyak-Ruppert averaging is a widely used technique to achieve the optimal asymptotic variance of stochastic approximation (SA) algorithms, yet its high-probability performance guarantees remain underexplored in general …

Q-Learning

Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs

2025-05-26 · Antoine Moulin, Gergely Neu, Luca Viano

We study the problem of offline imitation learning in Markov decision processes (MDPs), where the goal is to learn a well-performing policy given a dataset of state-action pairs generated by an expert policy. Complementi…

Imitation LearningQ-Learning

Distributionally Robust Deep Q-Learning

2025-05-25 · Chung I Lu, Julian Sester, Aijia Zhang

We propose a novel distributionally robust $Q$-learning algorithm for the non-tabular case accounting for continuous state spaces where the state transition of the underlying Markov decision process is subject to model u…

Q-Learning
1–20 / 1,918 다음 →