DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning
This paper introduces DW-Bench, a new benchmark that evaluates large language models (LLMs) on graph-topology reasoning over data warehouse schemas, explicitly integrating both foreign-key (FK) and data-lineage edges. The benchmark comprises 1,046 automatically generated, verifiably correct questions across five schemas. Experiments show that tool-augmented methods substantially outperform static approaches but plateau on hard compositional subtypes.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
From Gridworlds to Warehouses: Adapting Lightweight One-shot Multi-Agent Pathfinding for AGVs
Multi-agent pathfinding (MAPF) under one-shot planning is a core component of warehouse automation, yet classical formulations typically assume four-connected 2D grids with unit-time moves in four directions. To fill rea…
TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
Benchmarking is the de-facto standard for evaluating LLMs, due to its speed, replicability and low cost. However, recent work has pointed out that the majority of the open source benchmarks available today have been cont…
BenchmarkingDeep Reinforcement Learning for Dynamic Order Picking in Warehouse Operations
Order picking is a pivotal operation in warehouses that directly impacts overall efficiency and profitability. This study addresses the dynamic order picking problem, a significant concern in modern warehouse management,…
BenchmarkingDeep Reinforcement LearningManagementreinforcement-learningQUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs
The rise of large language models (LLMs) has created a need for advanced benchmarking systems beyond traditional setups. To this end, we introduce QUENCH, a novel text-based English Quizzing Benchmark manually curated an…
BenchmarkingCommon Sense ReasoningWorld KnowledgeGraphArena: Benchmarking Large Language Models on Graph Computational Problems
The "arms race" of Large Language Models (LLMs) demands novel, challenging, and diverse benchmarks to faithfully examine their progresses. We introduce GraphArena, a benchmarking tool designed to evaluate LLMs on graph c…
BenchmarkingHallucinationKnowledge Graphs