paper-with-me

홈 › Papers

Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models

2025-05-26 · Lachlan McGinness, Peter Baumgartner

Empirical methods to examine the capability of Large Language Models (LLMs) to use Automated Theorem Prover (ATP) reasoning strategies are studied. We evaluate the performance of State of the Art models from December 2023 and August 2024 on PRONTOQA steamroller reasoning problems. For that, we develop methods for assessing LLM response accuracy and correct answer correlation. Our results show that progress in improving LLM reasoning abilities has stalled over the nine month period. By tracking completion tokens, we show that almost all improvement in reasoning ability since GPT-4 was released can be attributed to either hidden system prompts or the training of models to automatically use generic Chain of Thought prompting strategies. Among the ATP reasoning strategies tried, we found that current frontier LLMs are best able to follow the bottom-up (also known as forward-chaining) strategy. A low positive correlation was found between an LLM response containing correct reasoning and arriving at the correct conclusion.

📄 PDF Abstract BibTeX arXiv:2505.19676

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Position-Wise Feed-Forward Layer 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Multi-Head Attention 설명 없음
Attention 설명 없음

Similar Papers 제목 키워드 기반

Complementarity Between Paid and Organic Installs in Mobile App Advertising

2025-04-22 · Harang Ju, Michael Zhao, Sinan Aral

Prior spending shutoff experiments in search advertising have found that paid ads cannibalize organic traffic. But it is unclear whether the same is true for other high volume advertising channels like mobile display adv…

Marketing

Dissertation on Applied Microeconomics of Freemium Pricing Strategies in Mobile App Market

2023-04-22 · Naixin Zhu

In my dissertation, I will analyze how the product market position of a mobile app affects its pricing strategies, which in turn impacts an app's monetization process. Using natural language processing and k-mean cluster…

Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs

2025-12-23 · Dhruv Anand, Ehsan Shareghi arxiv

We introduce Cube Bench, a Rubik's-cube benchmark for evaluating spatial and sequential reasoning in multimodal large language models (MLLMs). The benchmark decomposes performance into five skills: (i) reconstructing cub…

Spatial ReasoningVisual Reasoning

Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs

2025-09-28 · Shulin Huang, Yiran Ding, Junshu Pan, Yue Zhang arxiv

Enhancing the complex reasoning capabilities of Large Language Models (LLMs) attracts widespread attention. While reinforcement learning (RL) has shown superior performance for improving complex reasoning, its impact on …

Reinforcement Learning

Base Models Know How to Reason, Thinking Models Learn When

2025-10-08 · Constantin Venhoff, Iván Arcuschin, Philip Torr, Arthur Conmy 외 arxiv

What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model's reasoning behaviors by training small Sparse Autoencoders on sentence-l…