paper-with-me

Papers

Certified-Gap Dual-Price Policies for Real-Time Truckload Bid Acceptance with Relocating, Clock-Constrained Resources

2026-07-18 · Aswin Chandrasekaran arxiv

A truckload carrier must accept or reject each load tender within seconds. The decision depends on fleet state, hours-of-service (HOS) clocks, and appointment windows. We model this as a weakly coupled dynamic program in which the resources relocate and carry clocks: serving a request moves the truck to a new market and depletes its clocks, and whether a truck can serve a request depends on its state. Occupancy-based reusable-resource models do not cover this setting. We build a real-time dual-price policy from the same Lagrangian relaxation that gives the problem's upper bound. Policy and bound come from one object, so every run reports a certified optimality gap. We prove three things. First, the certificate is valid for any duals, any discretization, and any surrogate quality. Second, the policy's same-time spatial-gradient rule is exactly fluid complementary slackness, and the policy is asymptotically optimal in the subcritical fluid regime; the fitted prices are also portable across sample paths, by linear-programming basis stability. Third, certificates have limits: per-resource Lagrangian slack can stay bounded away from zero at every fleet size. We exhibit a three-truck kernel with an exact rational certificate and a replication lemma. On a public closed-loop benchmark with thirty paired seeds, the policy -- which needs no rollout labels, only one offline dual solve -- beats a rollout-trained surrogate on two of three scenarios (tight: +2.0 pp, 95% CI [+0.5, +3.6], Wilcoxon p = 0.023; mild: +3.5 pp, CI [+2.4, +4.5]) and ties the third. It decides in 0.04-0.09 ms, three orders of magnitude faster than the Monte Carlo rollout teacher. Its certificates are stable across ten bounded instances per scenario, at 57-64% of optimal, within 3-6 points of what the 1000x-slower teacher certifies.

📄 PDF Abstract BibTeX arXiv:2607.16891

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Support-aware offline policy selection for advertising marketplaces

2026-05-20 · Prashant Shekhar, Caroline Howard arxiv

Logged advertising auctions make offline reserve-price evaluation attractive but risky. Replay tables can identify policies with large apparent yield gains, yet they can also hide weak threshold support, multiple-compari…

CSymPlan: Certified Symbolic Planning and Control for High-DOF Manipulators

2026-08-24 · Aditya Narendra, Ashok Kumar Saini, Mahathi Anand, Mahmoud Khaled 외 arxiv

Robot manipulators are commonly engineered around a decoupled motion-generation stack: a planner computes a collision-free path and a lower-level controller tracks the resulting reference. This separation is computationa…

Price of Fairness in Short-Term and Long-Term Algorithmic Selections

2026-05-07 · Shahin Jabbari, Chen Wang arxiv

Algorithmic decision-making in high-stakes settings can have profound impacts on individuals and populations. While much prior work studies fairness in static settings, recent results show that enforcing static fairness …

CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models

2026-08-21 · Hui Lu, Zhijie Peng, Yuqi Lin, Zaijia Yang 외 arxiv

Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target discrete labels and cannot directly certify continuous, temporally correlated actions…

Pruning Cannot Hurt Robustness: Certified Trade-offs in Reinforcement Learning

2025-10-14 · James Pedley, Benjamin Etheridge, Stephen J. Roberts, Francesco Quinzan arxiv

Reinforcement learning (RL) policies deployed in real-world environments must remain reliable under adversarial perturbations. At the same time, modern deep RL agents are heavily over-parameterized, raising costs and fra…

Reinforcement Learning