paper-with-me

홈 › Papers

The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution

2025-08-17 · Elon Ezra, Ariel Weizman, Amos Azaria arxiv

Large language models (LLMs) are commonly evaluated on tasks that test their knowledge or reasoning abilities. In this paper, we explore a different type of evaluation: whether an LLM can predict aspects of its own responses. Since LLMs lack the ability to execute themselves, we introduce the Self-Execution Benchmark, which measures a model's ability to anticipate properties of its output, such as whether a question will be difficult for it, whether it will refuse to answer, or what kinds of associations it is likely to produce. Our experiments show that models generally perform poorly on this benchmark, and that increased model size or capability does not consistently lead to better performance. These results suggest a fundamental limitation in how LLMs represent and reason about their own behavior.

📄 PDF Abstract BibTeX arXiv:2508.12277

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs

2025-09-11 · Akshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab 외 arxiv

Does continued scaling of large language models (LLMs) yield diminishing returns? In this work, we show that short-task benchmarks may give an illusion of slowing progress, as even marginal gains in single-step accuracy …

CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation

2023-11-14 · Weixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li 외

Large Language Models (LLMs) have demonstrated remarkable performance on assisting humans in programming and facilitating programming automation. However, existing benchmarks for evaluating the code understanding and gen…

Code Generation

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models

2025-01-27 · Sudarshan Kamath Barkur, Sigurd Schacht, Johannes Scholl

Recent advances in Large Language Models (LLMs) have incorporated planning and reasoning capabilities, enabling models to outline steps before execution and provide transparent reasoning paths. This enhancement has reduc…

Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility

2026-01-19 · Nickil Maveli, Antonio Vergari, Shay B. Cohen arxiv

LLMs demonstrate strong performance on code benchmarks, yet consistent reasoning across forward and backward execution remains elusive. We present RoundTripCodeEval (RTCE), a benchmark of four code execution reasoning ta…

Self-Execution Simulation Improves Coding Models

2026-03-11 · Gallil Maimon, Ori Yoran, Felix Kreuk, Michael Hassid 외 arxiv

A promising research direction in enabling LLMs to generate consistently correct code involves addressing their inability to properly estimate program execution, particularly for code they generate. In this work, we demo…

Reinforcement Learning