paper-with-me

Papers

Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning

2025-05-21 · Tiasa Singha Roy, Aditeya Baral, Ayush Rajesh Jhaveri, Yusuf Baig

Large language models (LLMs) demonstrate considerable potential in various natural language tasks but face significant challenges in mathematical reasoning, particularly in executing precise, multi-step logic. However, current evaluation frameworks judge their performance solely based on accuracy, which only accounts for the final answer. This study explores these pitfalls by employing a novel evaluation framework. We propose an evaluation metric called the MAPLE score, which holistically quantifies reasoning misalignment by integrating error rates, redundancy, and validity.

📄 PDF Abstract BibTeX arXiv:2505.15623

Code (0)

등록된 구현이 없습니다.

Tasks

MathMathematical Reasoning

Similar Papers 제목 키워드 기반

Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems

2024-04-23 · Qihuang Zhong, Kang Wang, Ziyang Xu, Juhua Liu 외

Chain-of-Thought (CoT) prompting has enhanced the performance of Large Language Models (LLMs) across various reasoning tasks. However, CoT still falls short in dealing with complex math word problems, as it usually suffe…

Arithmetic ReasoningGSM8KMathMath Word Problem Solving

Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs

2024-05-24 · Siyuan Guo, Aniket Didolkar, Nan Rosemary Ke, Anirudh Goyal 외

We are beginning to see progress in language model assisted scientific discovery. Motivated by the use of LLMs as a general scientific assistant, this paper assesses the domain knowledge of LLMs through its understanding…

In-Context LearningLanguage ModelingLanguage ModellingMath+1

ConvexBench: Can LLMs Recognize Convex Functions?

2026-02-01 · Yepeng Liu, Yu Huang, Yu-Xiang Wang, Yingbin Liang 외 arxiv

Convex analysis is a modern branch of mathematics with many applications. As Large Language Models (LLMs) start to automate research-level math and sciences, it is important for LLMs to demonstrate the ability to underst…

Large language models in medicine: the potentials and pitfalls

2023-08-31 · Jesutofunmi A. Omiye, Haiwen Gui, Shawheen J. Rezaei, James Zou 외

Large language models (LLMs) have been applied to tasks in healthcare, ranging from medical exam questions to responding to patient questions. With increasing institutional partnerships between companies producing LLMs a…

Creating an AI Observer: Generative Semantic Workspaces

2024-06-07 · Pavan Holur, Shreyas Rajesh, David Chong, Vwani Roychowdhury

An experienced human Observer reading a document -- such as a crime report -- creates a succinct plot-like $\textit{``Working Memory''}$ comprising different actors, their prototypical roles and states at any point, thei…

Sentence