paper-with-me

홈 › Papers

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

2026-05-28 · Shahinul Hoque, Jinghuai Zhang, Jinyuan Sun, Fnu Suya arxiv

Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affects what users pay. We show that this kind of billing is hard to audit by design: providers hide the model, the tokenizer, and the execution to protect their IP, mitigate jailbreaks, and preserve user privacy, which means an auditor can only inspect proofs the provider supplies. The audit therefore reduces to a consistency check on the provider's own reports. We call this a trust paradox: every audit must trust some artifact, but current frameworks trust exactly the ones a provider has the strongest reason to manipulate. We study three recent token auditing frameworks and show that a provider with ordinary commercial capabilities can systematically inflate billed token counts. In the most permissive setting, hidden reasoning usage can be inflated by 1,469% on average without detection. At current frontier reasoning prices, that turns a \$100 honest bill into roughly a \$1,569 bill on the same query. Even when the user can see the full reasoning string, tokenization ambiguity alone still allows 50.85% over-reporting below the detection threshold. These results suggest the problem is not in any specific auditor but in any audit whose evidence comes from the audited party. Restoring honest billing will require verification that ties reported token counts to evidence the provider does not control, such as trusted execution attestation, cryptographic proofs of inference, or third-party re-execution.

📄 PDF Abstract BibTeX arXiv:2605.30040

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pay for The Second-Best Service: A Game-Theoretic Approach Against Dishonest LLM Providers

2025-11-02 · Yuhan Cao, Yu Wang, Sitong Liu, Miao Li 외 arxiv

The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability: the potential for dishonest manipulation by service providers. This manipulation…

Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives

2025-05-27 · Ander Artola Velasco, Stratis Tsirtsis, Nastaran Okati, Manuel Gomez-Rodriguez

State-of-the-art large language models require specialized hardware and substantial energy to operate. As a consequence, cloud-based services that provide access to large language models have become very popular. In thes…

Chatbot

CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs

2025-05-19 · Guoheng Sun, Ziyao Wang, Bowei Tian, Meng Liu 외

As post-training techniques evolve, large language models (LLMs) are increasingly augmented with structured multi-step reasoning abilities, often optimized through reinforcement learning. These reasoning-enhanced models …

Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation

2025-07-29 · Ziyao Wang, Guoheng Sun, Yexiao He, Zheyu Shen 외 arxiv

Commercial LLM services often conceal internal reasoning traces while still charging users for every generated token, including those from hidden intermediate steps, raising concerns of token inflation and potential over…

Verification of Machine Unlearning is Fragile

2024-08-01 · Binchi Zhang, Zihan Chen, Cong Shen, Jundong Li

As privacy concerns escalate in the realm of machine learning, data owners now have the option to utilize machine unlearning to remove their data from machine learning models, following recent legislation. To enhance tra…

Machine Unlearning