NeFut Logo NeFut
Admin Login

[CS.AI] What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets

Published at: 2026-09-10 22:00 Last updated: 2026-09-12 06:35
#algorithm #Machine Learning #LLM

We present a continuous, population‑scale measurement record covering two systems that share a single design lineage of autonomous language‑model trading agents in production. The first system, DX Terminal Pro, comprises 3,505 user‑funded vaults trading real ETH in Base memecoin markets for 21 days (Feb‑Mar 2026). The second, the DXAP live alpha fleet, contains 500‑599 user‑created agents with full history, 91‑117 concurrently active, trading Hyperliquid perpetuals (Jun‑Aug 2026). The record spans roughly six months, 7.5 M single‑model invocations, about 300 K on‑chain actions, and 231,638 multi‑tool turns producing 14,596 fills.

Four key findings emerge. First, the operating layer determines behavior more than any written strategy: a risk slider adds $+0.425$ leverage per level, agent fixed effects absorb ~60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75× at the top‑3 cut). Second, sizing is volatility‑blind: median leverage is $5.0\times$ in every volatility sextile, and a single posture‑slider cell (11% of the order book) accounts for 62% of liquidations. Third, agents capture almost none of the upside they encounter: 43.2% of positions see at least +300 bps of favorable excursion within 24 h, yet 49.3% of those close with negative returns; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is unprofitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% round‑trip win rate). A paired‑replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. All headlines survive day‑clustered inference, permutation nulls, and a common‑fee restatement; the paper closes with a 17‑rule methodological canon and our own retraction notice.

Review: This study delivers rare large‑scale empirical evidence, exposing the behavioral patterns and limitations of LLM trading agents in real markets and offering valuable guidance for future model refinement and risk management.

Original Source: https://arxiv.org/abs/2609.05663

[h] Back to Home