paper-with-me

Papers

Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

2025-02-20 · Axel Backlund, Lukas Petersson

While Large Language Models (LLMs) can exhibit impressive proficiency in isolated, short-term tasks, they often fail to maintain coherent performance over longer time horizons. In this paper, we present Vending-Bench, a simulated environment designed to specifically test an LLM-based agent's ability to manage a straightforward, long-running business scenario: operating a vending machine. Agents must balance inventories, place orders, set prices, and handle daily fees - tasks that are each simple but collectively, over long horizons (>20M tokens per run) stress an LLM's capacity for sustained, coherent decision-making. Our experiments reveal high variance in performance across multiple LLMs: Claude 3.5 Sonnet and o3-mini manage the machine well in most runs and turn a profit, but all models have runs that derail, either through misinterpreting delivery schedules, forgetting orders, or descending into tangential "meltdown" loops from which they rarely recover. We find no clear correlation between failures and the point at which the model's context window becomes full, suggesting that these breakdowns do not stem from memory limits. Apart from highlighting the high variance in performance over long time horizons, Vending-Bench also tests models' ability to acquire capital, a necessity in many hypothetical dangerous AI scenarios. We hope the benchmark can help in preparing for the advent of stronger AI systems.

📄 PDF Abstract BibTeX arXiv:2502.15840

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies

2026-02-10 · Xavier Hu, Jinxiang Xia, Shengze Xu, Kangqi Song 외 arxiv

Long-horizon planning is widely recognized as a core capability of autonomous LLM-based agents; however, current evaluation frameworks suffer from being largely episodic, domain-specific, or insufficiently grounded in pe…

Decision Making

AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations

2025-09-30 · Berdymyrat Ovezmyradov arxiv

The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent trends in AI benchmarking is performance of Large Language Models (LLMs) o…

Decision Making

Distributed Data Vending on Blockchain

2018-03-15 · Jiayu Zhou, Fengyi Tang, He Zhu, Ning Nan 외

Recent advances in blockchain technologies have provided exciting opportunities for decentralized applications. Specifically, blockchain-based smart contracts enable credible transactions without authorized third parties…

Retrieval

Predictive Maintenance Optimization for Smart Vending Machines Using IoT and Machine Learning

2025-06-26 · Md. Nisharul Hasan

The increasing proliferation of vending machines in public and commercial environments has placed a growing emphasis on operational efficiency and customer satisfaction. Traditional maintenance approaches either reactive…

Fault DetectionScheduling

Half-empty or half-full? A Hybrid Approach to Predict Recycling Behavior of Consumers to Increase Reverse Vending Machine Uptime

2020-03-30 · Jannis Walk, Robin Hirt, Niklas Kühl, Erik R. Hersløv

Reverse Vending Machines (RVMs) are a proven instrument for facilitating closed-loop plastic packaging recycling. A good customer experience at the RVM is crucial for a further proliferation of this technology. Bin full …

Time SeriesTime Series Analysis