paper-with-me

홈 › Papers

Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

2026-05-01 · Yuxuan Gao, Megan Wang, Yi Ling Yu arxiv

Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) tuple at which a specific quantization, decoding strategy, region, and serving stack is exposed. We introduce TokenArena, a continuous benchmark that measures inference at endpoint granularity along five core axes (output speed, time to first token, workload-blended price, effective context, and quality on the live endpoint) and synthesizes them, together with a modeled energy estimate, into three headline composites: joules per correct answer, dollars per correct answer, and endpoint fidelity (output-distribution similarity to a first-party reference). The framework's novelty is empirical and methodological. Across 78 endpoints serving 12 model families, the same model on different endpoints differs in mean accuracy by up to 12.5 points on math and code, in fingerprint similarity to first party by up to 12 points, in tail latency by an order of magnitude, and in modeled joules per correct answer by a factor of 6.2. We further show that workload-aware blended pricing reorders the leaderboard substantially: 7 of 10 top-ranked endpoints under the chat preset (3:1 input:output) fall out of the top 10 under the retrieval-augmented preset (20:1), and the reasoning preset (1:5) elevates frontier closed models that the chat preset penalizes on price. We release the framework, schema, probe and eval harness, and a v1.0 leaderboard snapshot under CC BY 4.0. TokenArena is a methodology, not a single ranking; we publish full provenance and limitations and welcome external replication.

📄 PDF Abstract BibTeX arXiv:2605.00300

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Energy-Arena: A Dynamic Benchmark for Operational Energy Forecasting

2026-04-27 · Max Kleinebrahm, Jonathan Berrisch, Philipp Eiser, Wolf Fichtner 외 arxiv

Energy forecasting research faces a persistent comparability gap that makes it difficult to measure consistent progress over time. Reported accuracy gains are often not directly comparable because models are evaluated un…

Time Series Forecasting

TS-Arena -- A Live Forecast Pre-Registration Platform

2025-12-23 · Marcel Meyer, Sascha Kaltenpoth, Henrik Albers, Kevin Zalipski 외 arxiv

Time Series Foundation Models (TSFMs) are transforming the field of forecasting. However, evaluating them on historical data is increasingly difficult due to the risks of train-test sample overlaps and temporal overlaps …

SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving

2025-05-29 · Wendong Xu, Jing Xiong, Chenyang Zhao, Qiujiang Chen 외

We present SwingArena, a competitive evaluation framework for Large Language Models (LLMs) that closely mirrors real-world software development workflows. Unlike traditional static benchmarks, SwingArena models the colla…

Code Generation

AMix-2: Establishing Protein as a Native Modality in Large Language Models

2026-05-29 · Keyue Qiu, Yixin Wu, Lihao Wang, Yawen Ouyang 외 arxiv

We present AMix-2, a protein-text foundation model that establishes protein as a native modality in large language models (LLMs), unifying protein understanding and sequence design within a single foundation model. AMix-…

The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations

2025-07-17 · Carlos Arriaga, Gonzalo Martínez, Eneko Sendin, Javier Conde 외

The evaluation of large language models is a complex task, in which several approaches have been proposed. The most common is the use of automated benchmarks in which LLMs have to answer multiple-choice questions of diff…

Language ModelingLanguage ModellingLarge Language ModelMultiple-choice