paper-with-me

홈 › Papers

How Inference Compute Shapes Frontier LLM Evaluation

2026-06-16 · Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, Cozmin Ududec arxiv

AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still report performance at a single restrictive budget, meaning that low scores may reflect the evaluation setup rather than the model's underlying capability. To test this, we evaluate up to 12 frontier language models on seven challenging benchmarks spanning software engineering, mathematics, medicine, and cybersecurity. We use a controlled setup combining three simple inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts, guided either by the model itself or by minimal correctness feedback. We find three main results. First, larger token budgets substantially improve performance on benchmarks across multiple domains, including cybersecurity, FrontierMath, Humanity's Last Exam, and TerminalBench. Second, fixed-budget evaluations can increasingly understate frontier capability as models advance. Newer models reach higher performance at large budgets, where they unlock harder tasks and solve them more reliably. Third, benchmarks differ in which inference-scaling methods help most: repeated submission broadly improves performance, but the value of larger token budgets, external feedback, and parallel attempts varies by benchmark. Overall, our results show that benchmark scores are protocol-dependent. We therefore argue that evaluations should report capability as a function of inference-time compute, specify protocol choices explicitly, and compare model generations over a large shared compute range at matched budgets, especially in safety- or policy-relevant settings.

📄 PDF Abstract BibTeX arXiv:2606.17930

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Inference Scaling Reshapes AI Governance

2025-02-12 · Toby Ord

The shift from scaling up the pre-training compute of AI systems to scaling up their inference compute may have profound effects on AI governance. The nature of these effects depends crucially on whether this new inferen…

PRISM: Pushing the Frontier of Deep Think via Process Reward Model-Guided Inference

2026-03-03 · Rituraj Sharma, Weiyuan Chen, Noah Provenzano, Tu Vu arxiv

DEEPTHINK methods improve reasoning by generating, refining, and aggregating populations of candidate solutions, which enables strong performance on complex mathematical and scientific tasks. However, existing frameworks…

The Inference-Compute Frontier and a Latency-Efficient Architecture for Limit Order Book Prediction

2026-06-24 · C. Evans Hedges arxiv

We study whether a scaling-law-style inference-compute frontier appears in limit order book prediction. Using FI-2010 and a suite of models ranging from small decision trees to neural LOB architectures, we find that the …

Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models

2025-12-31 · Ákos Prucs, Nara Csutora, Mátyás Antal, Márk Marosi arxiv

Large Language Models (LLMs) are demonstrating rapid improvements on complex reasoning benchmarks, particularly when allowed to utilize intermediate reasoning steps before converging on a final solution. However, current…

3D Optimization for AI Inference Scaling: Balancing Accuracy, Cost, and Latency

2025-10-21 · Minseok Jung, Abhas Ricky, Muhammad Rameez Chatni arxiv

AI inference scaling is often tuned through 1D heuristics (a fixed reasoning pass) or 2D bivariate trade-offs (e.g., accuracy vs. compute), which fail to consider cost and latency constraints. We introduce a 3D optimizat…