paper-with-me

Papers

Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

2024-04-02 · Philipp Mondorf, Barbara Plank

Large language models (LLMs) have recently shown impressive performance on tasks involving reasoning, leading to a lively debate on whether these models possess reasoning capabilities similar to humans. However, despite these successes, the depth of LLMs' reasoning abilities remains uncertain. This uncertainty partly stems from the predominant focus on task performance, measured through shallow accuracy metrics, rather than a thorough investigation of the models' reasoning behavior. This paper seeks to address this gap by providing a comprehensive review of studies that go beyond task accuracy, offering deeper insights into the models' reasoning processes. Furthermore, we survey prevalent methodologies to evaluate the reasoning behavior of LLMs, emphasizing current trends and efforts towards more nuanced reasoning analyses. Our review suggests that LLMs tend to rely on surface-level patterns and correlations in their training data, rather than on sophisticated reasoning abilities. Additionally, we identify the need for further research that delineates the key differences between human and LLM-based reasoning. Through this survey, we aim to shed light on the complex reasoning processes within LLMs.

📄 PDF Abstract BibTeX arXiv:2404.01869

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

2025-05-26 · Baihui Zheng, Boren Zheng, Kerui Cao, Yingshui Tan 외

Despite the remarkable proficiency of \textit{Large Reasoning Models} (LRMs) in handling complex reasoning tasks, their reliability in safety-critical scenarios remains uncertain. Existing evaluations primarily assess re…

Safety Alignment

Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning

2026-05-10 · Xia Yang, Xuanyi Zhang, Hao Hu, Feng Ji arxiv

Large language models now achieve high final-answer accuracy on mathematical reasoning benchmarks, but accuracy alone does not capture reasoning flexibility. We introduce a strategy-level evaluation framework instantiate…

Mathematical Reasoning

Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images

2026-01-14 · Xuchen Li, Xuzhao Li, Renjie Pi, Shiyu Hu 외 arxiv

Despite the remarkable progress of Vision-Language Models (VLMs) in adopting "Thinking-with-Images" capabilities, accurately evaluating the authenticity of their reasoning process remains a critical challenge. Existing b…

Visual Reasoning

Beyond Compliance: A Resistance-Informed Motivation Reasoning Framework for Challenging Psychological Client Simulation

2026-04-12 · Danni Liu, Bo Liu, Yuxin Hu, Hantao Zhao 외 arxiv

Psychological client simulators have emerged as a scalable solution for training and evaluating counselor trainees and psychological LLMs. Yet existing simulators exhibit unrealistic over-compliance, leaving counselors u…

Reinforcement LearningResponse Generation

Evaluating Mathematical Reasoning Beyond Accuracy

2024-04-08 · Shijie Xia, Xuefeng Li, Yixin Liu, Tongshuang Wu 외

The leaderboard of Large Language Models (LLMs) in mathematical tasks has been continuously updated. However, the majority of evaluations focus solely on the final results, neglecting the quality of the intermediate step…

MathMathematical Reasoning