paper-with-me

홈 › Papers

On the logical skills of large language models: evaluations using arbitrarily complex first-order logic problems

2025-02-20 · Shokhrukh Ibragimov, Arnulf Jentzen, Benno Kuckuck

We present a method of generating first-order logic statements whose complexity can be controlled along multiple dimensions. We use this method to automatically create several datasets consisting of questions asking for the truth or falsity of first-order logic statements in Zermelo-Fraenkel set theory. While the resolution of these questions does not require any knowledge beyond basic notation of first-order logic and set theory, it does require a degree of planning and logical reasoning, which can be controlled up to arbitrarily high difficulty by the complexity of the generated statements. Furthermore, we do extensive evaluations of the performance of various large language models, including recent models such as DeepSeek-R1 and OpenAI's o3-mini, on these datasets. All of the datasets along with the code used for generating them, as well as all data from the evaluations is publicly available at https://github.com/bkuckuck/logical-skills-of-llms.

📄 PDF Abstract BibTeX arXiv:2502.14180

Code (1)

bkuckuck/logical-skills-of-llms 공식 구현

Tasks

Logical Reasoning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models

2026-02-06 · Brian Rabern, Philipp Mondorf, Barbara Plank arxiv

Large language models perform well on many logical reasoning benchmarks, but it remains unclear which core logical skills they truly master. To address this, we introduce LogicSkills, a benchmark that isolates three fund…

Logical Reasoning

OPT-R: Exploring the Role of Explanations in Finetuning and Prompting for Reasoning Skills of Large Language Models

2023-05-19 · Badr AlKhamissi, Siddharth Verma, Ping Yu, Zhijing Jin 외

In this paper, we conduct a thorough investigation into the reasoning capabilities of Large Language Models (LLMs), focusing specifically on the Open Pretrained Transformers (OPT) models as a representative of such model…

DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models

2025-09-19 · Tsz Ting Chung, Lemao Liu, Mo Yu, Dit-Yan Yeung arxiv

Logic reasoning in natural language has been recognized as an important measure of human intelligence for Large Language Models (LLMs). Popular benchmarks may entangle multiple reasoning skills and thus provide unfaithfu…

Logical Reasoning

LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models

2024-01-01 · Yuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan 외

We introduce LogicAsker, a novel approach for evaluating and enhancing the logical reasoning capabilities of large language models (LLMs) such as ChatGPT and GPT-4. Despite LLMs' prowess in tasks like writing assistance,…

Code GenerationIn-Context LearningLogical ReasoningMachine Translation

Chat to Chip: Large Language Model Based Design of Arbitrarily Shaped Metasurfaces

2025-09-29 · Huanshu Zhang, Lei Kang, Sawyer D. Campbell, Douglas H. Werner arxiv

Traditional metasurface design is limited by the computational cost of full-wave simulations, preventing thorough exploration of complex configurations. Data-driven approaches have emerged as a solution to this bottlenec…