paper-with-me

Papers

Economic Evaluation of LLMs

2025-07-04 · Michael J. Zellinger, Matt Thomson arxiv

Practitioners often navigate LLM performance trade-offs by plotting Pareto frontiers of optimal accuracy-cost trade-offs. However, this approach offers no way to compare between LLMs with distinct strengths and weaknesses: for example, a cheap, error-prone model vs a pricey but accurate one. To address this gap, we propose economic evaluation of LLMs. Our framework quantifies the performance trade-off of an LLM as a single number based on the economic constraints of a concrete use case, all expressed in dollars: the cost of making a mistake, the cost of incremental latency, and the cost of abstaining from a query. We apply our economic evaluation framework to compare the performance of reasoning and non-reasoning models on difficult questions from the MATH benchmark, discovering that reasoning models offer better accuracy-cost tradeoffs as soon as the economic cost of a mistake exceeds \$0.01. In addition, we find that single large LLMs often outperform cascades when the cost of making a mistake is as low as \$0.1. Overall, our findings suggest that when automating meaningful human tasks with AI models, practitioners should typically use the most powerful available model, rather than attempt to minimize AI deployment costs, since deployment costs are likely dwarfed by the economic impact of AI errors.

📄 PDF Abstract BibTeX arXiv:2507.03834

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EconNLI: Evaluating Large Language Models on Economics Reasoning

2024-07-01 · Yue Guo, Yi Yang

Large Language Models (LLMs) are widely used for writing economic analysis reports or providing financial advice, but their ability to understand economic knowledge and reason about potential results of specific economic…

Decision MakingNatural Language Inference

Ideological Bias in LLMs' Economic Causal Reasoning

2026-04-23 · Donggyu Lee, Hyeok Yun, Jungwon Kim, Junsik Min 외 arxiv

Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly used in policy analysis and economic reporting, where directionally correct causa…

The Memorization Problem: Can We Trust LLMs' Economic Forecasts?

2025-04-20 · Alejandro Lopez-Lira, Yuehua Tang, Mingyin Zhu

Large language models (LLMs) cannot be trusted for economic forecasts during periods covered by their training data. We provide the first systematic evaluation of LLMs' memorization of economic and financial data, includ…

Memorization

Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition

2026-04-07 · Yushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu 외 arxiv

The ability of large language models (LLMs) to manage and acquire economic resources remains unclear. In this paper, we introduce \textbf{Market-Bench}, a comprehensive benchmark that evaluates the capabilities of LLMs i…

Macroeconomic Forecasting with Large Language Models

2024-07-01 · Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar

This paper presents a comparative analysis evaluating the accuracy of Large Language Models (LLMs) against traditional macro time series forecasting approaches. In recent times, LLMs have surged in popularity for forecas…

Time SeriesTime Series Forecasting