paper-with-me

홈 › Papers

Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning

2025-11-13 · Jason Chan, Zhixue Zhao, Robert Gaizauskas arxiv

Existing work investigates the reasoning capabilities of large language models (LLMs) to uncover their limitations, human-like biases and underlying processes. Such studies include evaluations of base LLMs (pre-trained on unlabeled corpora only) for this purpose. Our position paper argues that evaluating base LLMs' reasoning capabilities raises inherent methodological concerns that are overlooked in such existing studies. We highlight the fundamental mismatch between base LLMs' pretraining objective and normative qualities, such as correctness, by which reasoning is assessed. In particular, we show how base LLMs generate logically valid or invalid conclusions as coincidental byproducts of conforming to purely linguistic patterns of statistical plausibility. This fundamental mismatch challenges the assumptions that (a) base LLMs' outputs can be assessed as their bona fide attempts at correct answers or conclusions; and (b) conclusions about base LLMs' reasoning can generalize to post-trained LLMs optimized for successful instruction-following. We call for a critical re-examination of existing work that relies implicitly on these assumptions, and for future work to account for these methodological pitfalls.

📄 PDF Abstract BibTeX arXiv:2511.10381

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On Evaluating Adversarial Robustness of Chest X-ray Classification: Pitfalls and Best Practices

2022-12-15 · Salah Ghamizi, Maxime Cordy, Michail Papadakis, Yves Le Traon

Vulnerability to adversarial attacks is a well-known weakness of Deep Neural Networks. While most of the studies focus on natural images with standardized benchmarks like ImageNet and CIFAR, little research has considere…

Adversarial RobustnessClassificationMedical DiagnosisX-ray Classification

On Evaluating Adversarial Robustness

2019-02-18 · Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel 외

Correctly evaluating defenses against adversarial examples has proven to be extremely difficult. Despite the significant amount of recent work attempting to design defenses that withstand adaptive attacks, few have succe…

Adversarial AttackAdversarial DefenseAdversarial Robustness

ReLE: A Scalable System and Structured Benchmark for Diagnosing Capability Anisotropy in Chinese LLMs

2026-01-24 · Rui Fang, Jian Li, Wei Chen, Bin Hu 외 arxiv

Large Language Models (LLMs) have achieved rapid progress in Chinese language understanding, yet accurately evaluating their capabilities remains challenged by benchmark saturation and prohibitive computational costs. Wh…

Navigating Pitfalls: Evaluating LLMs in Machine Learning Programming Education

2025-05-23 · Smitha Kumar, Michael A. Lones, Manuel Maarek, Hind Zantout

The rapid advancement of Large Language Models (LLMs) has opened new avenues in education. This study examines the use of LLMs in supporting learning in machine learning education; in particular, it focuses on the abilit…

Model Selection

Generalizability of Machine Learning Models: Quantitative Evaluation of Three Methodological Pitfalls

2022-02-01 · Farhad Maleki, Katie Ovens, Rajiv Gupta, Caroline Reinhold 외

Purpose: Despite the potential of machine learning models, the lack of generalizability has hindered their widespread adoption in clinical practice. We investigate three methodological pitfalls: (1) violation of independ…

BIG-bench Machine LearningData Augmentationfeature selectionPneumonia Detection