paper-with-me

Papers

Offline Reinforcement Learning for Warehouse SLAM Throughput Control

2026-06-22 · Tina Dongxu Li, Mouhacine Benosman, Rajat Kumar, Kevin Tan, Ken Meszaros, Trevor Dardik arxiv

We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment. SLAM (Scan/Label/Apply/Manifest) throughput directly influences system congestion and operational efficiency. Our RL-based control approach dynamically recommends SLAM throughput settings that adaptively balance throughput maximization with downstream stability through intelligent adjustment of throttling behavior. We include a history-informed state representation, action space abstraction for delayed-impact control, and a reward function that captures both upstream and downstream operational metrics. Our approach is algorithm-agnostic, enabling integration of multiple offline RL methods under a unified architecture. We instantiate our framework with three state-of-the-art offline RL algorithms, and trained the models offline using de-identified historical operational logs from a large-scale warehouse. Policy performance is evaluated using a comprehensive multi-method strategy. These include model-free approaches including immediate reward estimation via regression models and long-horizon Fitted Q Evaluation (FQE), as well as model-based Deep Koopman dynamics evaluation. Empirical results reveal that the CQL policy consistently outperforms alternatives, improving system health by 22.97% and reducing average throttling duration by 3.18%. These findings demonstrate the potential of offline RL for safe and scalable warehouse throughput control optimization.

📄 PDF Abstract BibTeX arXiv:2606.23978

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningOffline RL

Similar Papers 제목 키워드 기반

Learning to Staff: Offline Reinforcement Learning and Fine-Tuned LLMs for Warehouse Staffing Optimization

2026-03-25 · Kalle Kujanpää, Yuying Zhu, Kristina Klinkner, Shervin Malmasi arxiv

We investigate machine learning approaches for optimizing real-time staffing decisions in semi-automated warehouse sortation systems. Operational decision-making can be supported at different levels of abstraction, with …

Reinforcement LearningOffline RL

Learning-guided Prioritized Planning for Lifelong Multi-Agent Path Finding in Warehouse Automation

2026-03-25 · Han Zheng, Yining Ma, Brandon Araki, Jingkai Chen 외 arxiv

Lifelong Multi-Agent Path Finding (MAPF) is critical for modern warehouse automation, which requires multiple robots to continuously navigate conflict-free paths to optimize the overall system throughput. However, the co…

Reinforcement Learning

GRAND: Guidance, Rebalancing, and Assignment for Networked Dispatch in Multi-Agent Path Finding

2025-12-02 · Johannes Gaber, Meshal Alharbi, Daniele Gammelli, Gioele Zardini arxiv

Large robot fleets are now common in warehouses and other logistics settings, where small control gains translate into large operational impacts. In this article, we address task scheduling for lifelong Multi-Agent Picku…

Reinforcement LearningGraph Neural Network

Multi-Robot Coordination and Layout Design for Automated Warehousing

2023-05-10 · Yulun Zhang, Matthew C. Fontaine, Varun Bhatt, Stefanos Nikolaidis 외

With the rapid progress in Multi-Agent Path Finding (MAPF), researchers have studied how MAPF algorithms can be deployed to coordinate hundreds of robots in large automated warehouses. While most works try to improve the…

DiversityLayout DesignMulti-Agent Path Finding

Stress-Relief Annealing: Polynomial-Time Simulation-Free Layout Optimization for Automated Warehouses

2026-08-02 · Xiangjie Luo, Yulun Zhang, Miyuki Koshimura, Makoto Yokoo 외 arxiv

We study the problem of optimizing physical layouts for automated warehouses, where hundreds to thousands of robots are coordinated to transport packages. Previous works have shown that optimizing the warehouse layout (e…