paper-with-me

Papers

Evaluating the Ability of Large Language Models to Reason about Cardinal Directions

2024-06-24 · Anthony G Cohn, Robert E Blackwell

We investigate the abilities of a representative set of Large language Models (LLMs) to reason about cardinal directions (CDs). To do so, we create two datasets: the first, co-created with ChatGPT, focuses largely on recall of world knowledge about CDs; the second is generated from a set of templates, comprehensively testing an LLM's ability to determine the correct CD given a particular scenario. The templates allow for a number of degrees of variation such as means of locomotion of the agent involved, and whether set in the first , second or third person. Even with a temperature setting of zero, Our experiments show that although LLMs are able to perform well in the simpler dataset, in the second more complex dataset no LLM is able to reliably determine the correct CD, even with a temperature setting of zero.

📄 PDF Abstract BibTeX arXiv:2406.16528

Code (0)

등록된 구현이 없습니다.

Tasks

World Knowledge

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited

2025-07-16 · Anthony G Cohn, Robert E Blackwell arxiv

We investigate the abilities of 28 Large language Models (LLMs) to reason about cardinal directions (CDs) using a benchmark generated from a set of templates, extensively testing an LLM's ability to determine the correct…

CommonWhy: A Dataset for Evaluating Entity-Based Causal Commonsense Reasoning in Large Language Models

2026-05-13 · Armin Toroghi, Faeze Moradi Kalarde, Scott Sanner arxiv

To effectively interact with the real world, Large Language Models (LLMs) require entity-based commonsense reasoning, a challenging task that necessitates integrating factual knowledge about specific entities with common…

Graph Question Answering

GPTEval: A Survey on Assessments of ChatGPT and GPT-4

2023-08-24 · Rui Mao, Guanyi Chen, Xulang Zhang, Frank Guerin 외

The emergence of ChatGPT has generated much speculation in the press about its potential to disrupt social and economic systems. Its astonishing language ability has aroused strong curiosity among scholars about its perf…

Survey

PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

2022-06-21 · NeurIPS 2023 11 · Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan 외

Generating plans of action, and reasoning about change have long been considered a core competence of intelligent agents. It is thus no surprise that evaluating the planning and reasoning capabilities of large language m…

Common Sense ReasoningDiversityWorld Knowledge

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

2025-06-10 · Polina Kirichenko, Mark Ibrahim, Kamalika Chaudhuri, Samuel J. Bell

For Large Language Models (LLMs) to be reliably deployed in both everyday and high-stakes domains, knowing when not to answer is equally critical as answering correctly. Real-world user queries, which can be underspecifi…

Math