paper-with-me

Papers

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi

2026-05-18 · Anthony G. Cohn, Robert E. Blackwell arxiv

We introduce an extensive qualitative spatial and temporal reasoning (QSTR) benchmark for evaluating large language models (LLMs). We pose questions concerning compositional reasoning (using composition tables, CT), converse relations, and conceptual neighbourhoods (CN) for QSTR calculi, Point Algebra (PA), Allen's Interval Algebra, Interval and Duration (INDU), Region Connection Calculus (RCC-5, RCC-8, and RCC-22), the nine intersection model, cardinal direction calculus, and STAR. The RCC-22 CN is published here for the first time. An extended benchmark systematically varies question presentation including prefix/infix, words/symbols/nonce terms and schematic descriptions for selected calculi. We report results for contemporary frontier models. All models tested perform better than guessing but none can consistently answer all questions correctly. Performance varies sharply by calculus, with PA being the most straightforward, and RCC-22 the most difficult. We release the benchmark, and our results under an open licence to facilitate further assessment of qualitative spatio/temporal reasoning in LLMs.

📄 PDF Abstract BibTeX arXiv:2605.18380

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cost of Reasoning in non-English Languages: A Case Study on Japanese

2026-07-11 · Yuu Jinnai arxiv

Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training data is most abundant. However, reasoning trace is a clue for model int…

FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

2023-11-16 · Yimin Jing, Renren Jin, Jiahao Hu, Huishi Qiu 외

The effective assessment of the instruction-following ability of large language models (LLMs) is of paramount importance. A model that cannot adhere to human instructions might be not able to provide reliable and helpful…

Instruction FollowingLogical ReasoningSpatial Reasoning

Evaluating Interventional Reasoning Capabilities of Large Language Models

2024-04-08 · Tejas Kasetty, Divyat Mahajan, Gintare Karolina Dziugaite, Alexandre Drouin 외

Numerous decision-making tasks require estimating causal effects under interventions on different parts of a system. As practitioners consider using large language models (LLMs) to automate decisions, studying their caus…

Causal InferenceDecision Making

Zero, Finite, and Infinite Belief History of Theory of Mind Reasoning in Large Language Models

2024-06-07 · Weizhi Tang, Vaishak Belle

Large Language Models (LLMs) have recently shown a promise and emergence of Theory of Mind (ToM) ability and even outperform humans in certain ToM tasks. To evaluate and extend the boundaries of the ToM reasoning ability…

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

2025-04-21 · Weiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen 외

Visual reasoning is a core component of human intelligence and a critical capability for advanced multimodal models. Yet current reasoning evaluations of multimodal large language models (MLLMs) often rely on text descri…

AttributeVisual Reasoning