paper-with-me

홈 › Papers

DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning

2026-04-21 · Ahmed G. A. H Ahmed, C. Okan Sakar arxiv

This paper introduces DW-Bench, a new benchmark that evaluates large language models (LLMs) on graph-topology reasoning over data warehouse schemas, explicitly integrating both foreign-key (FK) and data-lineage edges. The benchmark comprises 1,046 automatically generated, verifiably correct questions across five schemas. Experiments show that tool-augmented methods substantially outperform static approaches but plateau on hard compositional subtypes.

📄 PDF Abstract BibTeX arXiv:2604.18964

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Gridworlds to Warehouses: Adapting Lightweight One-shot Multi-Agent Pathfinding for AGVs

2026-05-15 · Hiroki Nagai, Keisuke Okumura arxiv

Multi-agent pathfinding (MAPF) under one-shot planning is a core component of warehouse automation, yet classical formulations typically assume four-connected 2D grids with unit-time moves in four directions. To fill rea…

TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs

2024-03-01 · Tanmay Rajore, Nishanth Chandran, Sunayana Sitaram, Divya Gupta 외

Benchmarking is the de-facto standard for evaluating LLMs, due to its speed, replicability and low cost. However, recent work has pointed out that the majority of the open source benchmarks available today have been cont…

Benchmarking

Deep Reinforcement Learning for Dynamic Order Picking in Warehouse Operations

2024-08-03 · Sasan Mahmoudinazlou, Abhay Sobhanan, Hadi Charkhgard, Ali Eshragh 외

Order picking is a pivotal operation in warehouses that directly impacts overall efficiency and profitability. This study addresses the dynamic order picking problem, a significant concern in modern warehouse management,…

BenchmarkingDeep Reinforcement LearningManagementreinforcement-learning

QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs

2024-12-16 · Mohammad Aflah Khan, Neemesh Yadav, Sarah Masud, Md. Shad Akhtar

The rise of large language models (LLMs) has created a need for advanced benchmarking systems beyond traditional setups. To this end, we introduce QUENCH, a novel text-based English Quizzing Benchmark manually curated an…

BenchmarkingCommon Sense ReasoningWorld Knowledge

GraphArena: Benchmarking Large Language Models on Graph Computational Problems

2024-06-29 · Jianheng Tang, Qifan Zhang, Yuhan Li, Jia Li

The "arms race" of Large Language Models (LLMs) demands novel, challenging, and diverse benchmarks to faithfully examine their progresses. We introduce GraphArena, a benchmarking tool designed to evaluate LLMs on graph c…

BenchmarkingHallucinationKnowledge Graphs