paper-with-me

홈 › Papers

Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology

2024-10-19 · Wei Xie, Shuoyoucheng Ma, Zhenhua Wang, Enze Wang, Kai Chen, Xiaobing Sun, Baosheng Wang

The cognitive mechanism by which Large Language Models (LLMs) solve mathematical problems remains a widely debated and unresolved issue. Currently, there is little interpretable experimental evidence that connects LLMs' problem-solving with human cognitive psychology.To determine if LLMs possess human-like mathematical reasoning, we modified the problems used in the human Cognitive Reflection Test (CRT). Our results show that, even with the use of Chains of Thought (CoT) prompts, mainstream LLMs, including the latest o1 model (noted for its reasoning capabilities), have a high error rate when solving these modified CRT problems. Specifically, the average accuracy rate dropped by up to 50% compared to the original questions.Further analysis of LLMs' incorrect answers suggests that they primarily rely on pattern matching from their training data, which aligns more with human intuition (System 1 thinking) rather than with human-like reasoning (System 2 thinking). This finding challenges the belief that LLMs have genuine mathematical reasoning abilities comparable to humans. As a result, this work may adjust overly optimistic views on LLMs' progress towards artificial general intelligence.

📄 PDF Abstract BibTeX arXiv:2410.14979

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningMathMathematical Reasoning

Similar Papers 제목 키워드 기반

Free-form language-based robotic reasoning and grasping

2025-03-17 · Runyu Jiao, Alice Fasoli, Francesco Giuliari, Matteo Bortolon 외

Performing robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spatial relationships between objects. Visio…

FormRobotic GraspingSpatial ReasoningWorld Knowledge

AI for Mathematics: A Cognitive Science Perspective

2023-10-19 · Cedegao E. Zhang, Katherine M. Collins, Adrian Weller, Joshua B. Tenenbaum

Mathematics is one of the most powerful conceptual systems developed and used by the human species. Dreams of automated mathematicians have a storied history in artificial intelligence (AI). Rapid progress in AI, particu…

ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models

2023-09-30 · Zhaowei Zhang, Fengshuo Bai, Jun Gao, Yaodong Yang

Personal values are a crucial factor behind human decision-making. Considering that Large Language Models (LLMs) have been shown to impact human decisions significantly, it is essential to make sure they accurately under…

Decision Making

Formal Mathematical Reasoning: A New Frontier in AI

2024-12-20 · Kaiyu Yang, Gabriel Poesia, Jingxuan He, Wenda Li 외

AI for Mathematics (AI4Math) is not only intriguing intellectually but also crucial for AI-driven discovery in science, engineering, and beyond. Extensive efforts on AI4Math have mirrored techniques in NLP, in particular…

Automated Theorem ProvingMathMathematical Reasoning

Unlearning as Ablation: Toward a Falsifiable Benchmark for Generative Scientific Discovery

2025-08-25 · Robert Yang arxiv

Bold claims about AI's role in science-from "AGI will cure all diseases" to promises of radically accelerated discovery-raise a central epistemic question: do large language models (LLMs) truly generate new knowledge, or…