paper-with-me

홈 › Papers

R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling

2025-08-21 · Raj Jain, Marc Wetter arxiv

Effective scheduling under tight resource, timing, and operational constraints underpins large-scale planning across sectors such as capital projects, manufacturing, logistics, and IT fleet transitions. However, the reliability of large language models (LLMs) when reasoning under high-constraint regimes is insufficiently characterized. To address this gap, we present R-ConstraintBench, a scalable framework that evaluates models on Resource-Constrained Project Scheduling Problems (RCPSP), an NP-Complete feasibility class, while difficulty increases via linear growth in constraints. R-ConstraintBench incrementally increases non-redundant precedence constraints in Directed Acyclic Graphs (DAGs) and then introduces downtime, temporal windows, and disjunctive constraints. As an illustrative example, we instantiate the benchmark in a data center migration setting and evaluate multiple LLMs using feasibility and error analysis, identifying degradation thresholds and constraint types most associated with failure. Empirically, strong models are near-ceiling on precedence-only DAGs, but feasibility performance collapses when downtime, temporal windows, and disjunctive constraints interact, implicating constraint interaction, not graph depth, as the principal bottleneck. Performance on clean synthetic ramps also does not guarantee transfer to domain-grounded scenarios, underscoring limited generalization.

📄 PDF Abstract BibTeX arXiv:2508.15204

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ConstraintBench: Benchmarking LLM Constraint Reasoning on Direct Optimization

2026-02-25 · Joseph Tso, Preston Schmittou, Quan Huynh, Jibran Hutchins arxiv

Large language models are increasingly applied to operational decision-making where the underlying structure is constrained optimization. Existing benchmarks evaluate whether LLMs can formulate optimization problems as s…

SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling

2026-08-02 · Shrenil Shaun Sharma, Avi Sharma arxiv

This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form variation. Grounded in canonical scheduling instances and solver-derived f…

TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks

2025-11-03 · Hanwen Xu, Xuyao Huang, Yuzhe Liu, Kai Yu 외 arxiv

Large language model (LLM) agents have exhibited strong problem-solving competence across domains like research and coding. Yet, it remains underexplored whether LLM agents can tackle compounding real-world problems that…

Reinforcement Learning

Agentic AI Home Energy Management System: A Large Language Model Framework for Residential Load Scheduling

2025-10-30 · Reda El Makroum, Sebastian Zwickl-Bernhard, Lukas Kranzl arxiv

The electricity sector transition requires substantial increases in residential demand response capacity, yet Home Energy Management Systems (HEMS) adoption remains limited by user interaction barriers requiring translat…

Prompt Engineering

Planning in Strawberry Fields: Evaluating and Improving the Planning and Scheduling Capabilities of LRM o1

2024-10-03 · Karthik Valmeekam, Kaya Stechly, Atharva Gundawar, Subbarao Kambhampati

The ability to plan a course of action that achieves a desired state of affairs has long been considered a core competence of intelligent agents and has been an integral part of AI research since its inception. With the …

Scheduling