paper-with-me

홈 › Papers

Learning to Solve Multiple-TSP with Time Window and Rejections via Deep Reinforcement Learning

2022-09-13 · Rongkai Zhang, Cong Zhang, Zhiguang Cao, Wen Song, Puay Siew Tan, Jie Zhang, Bihan Wen, Justin Dauwels

We propose a manager-worker framework based on deep reinforcement learning to tackle a hard yet nontrivial variant of Travelling Salesman Problem (TSP), \ie~multiple-vehicle TSP with time window and rejections (mTSPTWR), where customers who cannot be served before the deadline are subject to rejections. Particularly, in the proposed framework, a manager agent learns to divide mTSPTWR into sub-routing tasks by assigning customers to each vehicle via a Graph Isomorphism Network (GIN) based policy network. A worker agent learns to solve sub-routing tasks by minimizing the cost in terms of both tour length and rejection rate for each vehicle, the maximum of which is then fed back to the manager agent to learn better assignments. Experimental results demonstrate that the proposed framework outperforms strong baselines in terms of higher solution quality and shorter computation time. More importantly, the trained agents also achieve competitive performance for solving unseen larger instances.

📄 PDF Abstract BibTeX arXiv:2209.06094

Code (1)

zcaicaros/manager-worker-mtsptwr 공식 구현 pytorch

Tasks

Deep Reinforcement Learning

Similar Papers 제목 키워드 기반

Demand Acceptance using Reinforcement Learning for Dynamic Vehicle Routing Problem with Emission Quota

2026-02-27 · Farid Najar, Dominique Barth, Yann Strozecki arxiv

This paper introduces and formalizes the Dynamic and Stochastic Vehicle Routing Problem with Emission Quota (DS-QVRP-RR), a novel routing problems that integrates dynamic demand acceptance and routing with a global emiss…

Reinforcement Learning

Multi-Vehicle Routing Problems with Soft Time Windows: A Multi-Agent Reinforcement Learning Approach

2020-02-13 · Ke Zhang, Meng Li, Zhengchao Zhang, Xi Lin 외

Multi-vehicle routing problem with soft time windows (MVRPSTW) is an indispensable constituent in urban logistics distribution systems. Over the past decade, numerous methods for MVRPSTW have been proposed, but most are …

Computational EfficiencyDecoderMulti-agent Reinforcement Learningreinforcement-learning+2

Self-Distilled Agentic Reinforcement Learning

2026-05-14 · Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang 외 arxiv

Reinforcement learning (RL) has emerged as a central paradigm for post-training LLM agents, yet its trajectory-level reward signal provides only coarse supervision for long-horizon interaction. On-Policy Self-Distillatio…

Reinforcement Learning

Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism

2026-01-09 · Yuhao Shen, Tianyu Liu, Junyi Shen, Jinyang Wu 외 arxiv

Parallel Speculative Decoding (PSD) accelerates traditional Speculative Decoding (SD) by overlapping draft generation with verification. However, it remains hampered by two fundamental challenges: (1) a theoretical speed…

Deep Reinforcement Learning for Electric Vehicle Routing Problem with Time Windows

2020-10-05 · Bo Lin, Bissan Ghaddar, Jatin Nathwani

The past decade has seen a rapid penetration of electric vehicles (EV) in the market, more and more logistics and transportation companies start to deploy EVs for service provision. In order to model the operations of a …

Deep Reinforcement LearningGraph Embeddingreinforcement-learningReinforcement Learning (RL)