paper-with-me

홈 › Papers

When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents

2025-10-13 · Lingfei Qian, Xueqing Peng, Yan Wang, Vincent Jim Zhang, Huan He, Hanley Smith, Yi Han, Yueru He, Haohang Li, Yupeng Cao, Yangyang Yu, Alejandro Lopez-Lira, Peng Lu, Jian-Yun Nie, Guojun Xiong, Jimin Huang, Sophia Ananiadou arxiv

Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies test models instead of agents, cover limited periods and assets, and rely on unverified data. To address these gaps, we introduce Agent Market Arena (AMA), the first lifelong, real-time benchmark for evaluating LLM-based trading agents across multiple markets. AMA integrates verified trading data, expert-checked news, and diverse agent architectures within a unified trading framework, enabling fair and continuous comparison under real conditions. It implements four agents, including InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning, and evaluates them across GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on both cryptocurrency and stock markets demonstrate that agent frameworks display markedly distinct behavioral patterns, spanning from aggressive risk-taking to conservative decision-making, whereas model backbones contribute less to outcome variation. AMA thus establishes a foundation for rigorous, reproducible, and continuously evolving evaluation of financial reasoning and trading intelligence in LLM-based agents.

📄 PDF Abstract BibTeX arXiv:2510.11695

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LiveTradeBench: Seeking Real-World Alpha with Large Language Models

2025-11-05 · Haofei Yu, Fenghai Li, Jiaxuan You arxiv

Large language models (LLMs) achieve strong performance across benchmarks--from knowledge quizzes and math reasoning to web-agent tasks--but these tests occur in static settings, lacking real dynamics and uncertainty. Co…

Decision Making

CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market

2026-06-30 · Yao Shi, Kingfung Luo, Nan Tang, Yuyu Luo arxiv

Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules. These properties make them hard for traditional quantitative models, …

Robust Reinforcement Learning in Finance: Modeling Market Impact with Elliptic Uncertainty Sets

2025-10-22 · Shaocong Ma, Heng Huang arxiv

In financial applications, reinforcement learning (RL) agents are commonly trained on historical data, where their actions do not influence prices. However, during deployment, these agents trade in live markets where the…

Reinforcement Learning

Emergence of Cooperative Long-term Market Loyalty in Double Auction Markets

2017-08-30

Loyal buyer-seller relationships can arise by design, e.g. when a seller tailors a product to a specific market niche to accomplish the best possible returns, and buyers respond to the dedicated efforts the seller makes …

Optimal Trading with Differing Trade Signals

2020-06-24 · Ryan Donnelly, Matthew Lorig

We consider the problem of maximizing portfolio value when an agent has a subjective view on asset value which differs from the traded market price. The agent's trades will have a price impact which affect the price at w…