paper-with-me

홈 › Papers

Steamroller Problems: An Evaluation of LLM Reasoning Capability with Automated Theorem Prover Strategies

2024-07-17 · Lachlan McGinness, Peter Baumgartner

This study presents the first examination of the ability of Large Language Models (LLMs) to follow reasoning strategies that are used to guide Automated Theorem Provers (ATPs). We evaluate the performance of GPT4, GPT3.5 Turbo and Google's recent Gemini model on problems from a steamroller domain. In addition to determining accuracy we make use of the Natural Language Processing library spaCy to explore new methods of investigating LLM's reasoning capabilities. This led to one alarming result, the low correlation between correct reasoning and correct answers for any of the tested models. We found that the models' performance when using the ATP reasoning strategies was comparable to one-shot chain of thought and observe that attention to uncertainty in the accuracy results is critical when drawing conclusions about model performance. Consistent with previous speculation we confirm that LLMs have a preference for, and are best able to follow, bottom up reasoning processes. However, the reasoning strategies can still be beneficial for deriving small and relevant sets of formulas for external processing by a trusted inference engine.

📄 PDF Abstract BibTeX arXiv:2407.20244

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Library 설명 없음

Similar Papers 제목 키워드 기반

Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models

2025-05-26 · Lachlan McGinness, Peter Baumgartner

Empirical methods to examine the capability of Large Language Models (LLMs) to use Automated Theorem Prover (ATP) reasoning strategies are studied. We evaluate the performance of State of the Art models from December 202…

Automated Theorem Provers Help Improve Large Language Model Reasoning

2024-08-07 · Lachlan McGinness, Peter Baumgartner

In this paper we demonstrate how logic programming systems and Automated first-order logic Theorem Provers (ATPs) can improve the accuracy of Large Language Models (LLMs) for logical reasoning tasks where the baseline pe…

Formal LogicLanguage ModelingLanguage ModellingLarge Language Model+2

GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts

2025-09-29 · Fan Yuan, Yuchen Yan, Yifan Jiang, Haoran Zhao 외 arxiv

Mathematical reasoning is a key capability for vision-language models (VLMs), yet current benchmarks mainly evaluate text-based or explicitly symbolic visual inputs. It remains unclear whether VLMs can reason mathematica…

Mathematical Reasoning

MM-THEBench: Do Reasoning MLLMs Think Reasonably?

2026-01-30 · Zhidian Huang, Zijun Yao, Ji Qi, Shangqing Tu 외 arxiv

Recent advances in multimodal large language models (MLLMs) mark a shift from non-thinking models to post-trained reasoning models capable of solving complex problems through thinking. However, whether such thinking miti…

ProcessBench: Identifying Process Errors in Mathematical Reasoning

2024-12-09 · Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin 외

As language models regularly make mistakes when solving math problems, automated identification of errors in the reasoning process becomes increasingly significant for their scalable oversight. In this paper, we introduc…

GSM8KMathMathematical Reasoning