paper-with-me

홈 › Papers

What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets

2026-09-04 · T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau hf

We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.

📄 PDF Abstract BibTeX arXiv:2609.05663

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deep Reinforcement Learning for Active High Frequency Trading

2021-01-18 · Antonio Briola, Jeremy Turiel, Riccardo Marcaccioli, Alvaro Cauderan 외

We introduce the first end-to-end Deep Reinforcement Learning (DRL) based framework for active high frequency trading in the stock market. We train DRL agents to trade one unit of Intel Corporation stock by employing the…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

Achilles, Neural Network to Predict the Gold Vs US Dollar Integration with Trading Bot for Automatic Trading

2024-10-29 · NA 2024 10 · Angel Varela1

Predicting the stock market is a big challenge for the machine learning world. It is known how difficult it is to have accurate and consistent predictions with ML models. Some architectures are able to capture the mov…

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

2026-07-07 · Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev 외 hf

We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience …

Commodities Trading through Deep Policy Gradient Methods

2023-08-10 · Jonas Hanetho

Algorithmic trading has gained attention due to its potential for generating superior returns. This paper investigates the effectiveness of deep reinforcement learning (DRL) methods in algorithmic commodities trading. It…

Algorithmic TradingDeep Reinforcement LearningPolicy Gradient MethodsTime Series

Learning the Market: Sentiment-Based Ensemble Trading Agents

2024-02-02 · Andrew Ye, James Xu, Vidyut Veedgav, Yi Wang 외

We propose and study the integration of sentiment analysis and deep reinforcement learning ensemble algorithms for stock trading by evaluating strategies capable of dynamically altering their active agent given the concu…

Deep Reinforcement LearningSentiment Analysis