paper-with-me

홈 › Papers

How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem

2026-05-07 · Chun Zheng, Lianlong Wu, Bingqian Li, Lvting Liu, Yi Zhou arxiv

Large Language Models (LLMs) have achieved great improvements in recent years. Nevertheless, it still remains unclear how good LLMs are for reasoning tasks, especially for long-chain ones. In this paper, we evaluate LLMs' performance on the simplest yet long-chain reasoning task, namely the Equivalence Class Problem (ECP), i.e., determining whether two variables are equal given a set of randomly generated equivalence relations. We consider both reasoning and non-reasoning representative LLMs over a large variety of problem instances, ranging over different numbers of variables, connectivity probabilities, prompts, and other factors. The experimental results show that non-reasoning LLMs fail ECP, while reasoning models are significantly better but still struggle to completely solve this problem. Interestingly, considering various connectivity probabilities with a fixed number of variables, we observe that, for non-reasoning models, the hardest problem instances coincide with the phase transition point of ln n/(n-1), suggesting the chaos of the problem; in contrast, for reasoning models, the hardest ones coincide with the biggest diameter, suggesting the reasoning difficulty of the problem.

📄 PDF Abstract BibTeX arXiv:2605.06882

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition

2025-02-10 · Guanghao Ye, Khiem Duc Pham, Xinzhi Zhang, Sivakanth Gopi 외

Recent AI advancements, such as OpenAI's new models, are transforming LLMs into LRMs (Large Reasoning Models) that perform reasoning during inference, taking extra time and compute for higher-quality outputs. We aim to u…

Math

Humans and LLMs Diverge on Probabilistic Inferences

2026-02-26 · Gaurav Kamath, Sreenath Madathil, Sebastian Schuster, Marie-Catherine de Marneffe 외 arxiv

Human reasoning often involves working over limited information to arrive at probabilistic conclusions. In its simplest form, this involves making an inference that is not strictly entailed by a premise, but rather only …

Quantifying the Necessity of Chain of Thought through Opaque Serial Depth

2026-03-10 · Jonah Brown-Cohen, David Lindner, Rohin Shah arxiv

Large language models (LLMs) tend to externalize their reasoning in their chain of thought, making the chain of thought a good target for monitoring. This is partially an inherent feature of the Transformer architecture:…

Large Language Models as Test Case Generators: Performance Evaluation and Enhancement

2024-04-20 · Kefan Li, Yuan Yuan

Code generation with Large Language Models (LLMs) has been extensively studied and achieved remarkable progress. As a complementary aspect to code generation, test case generation is of crucial importance in ensuring the…

Code GenerationTest Case Creation

Large Language Models are few(1)-shot Table Reasoners

2022-10-13 · Wenhu Chen

Recent literature has shown that large language models (LLMs) are generally excellent few-shot reasoners to solve text reasoning tasks. However, the capability of LLMs on table reasoning tasks is yet to be explored. In t…

Fact VerificationIn-Context Learning