Source-linked AI summary
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
TL;DR
The paper asks what autonomous language-model trading agents actually do in production, beyond short-horizon leaderboard P&L. It assembles a six-month record across two production fleets, then finds operating-layer effects, volatility-blind risk, weak directional performance, and substantial limits in the study’s paper-trading mechanics.
Problem
Existing evaluations largely measure prompted models by leaderboard P&L over hours to days, leaving limited population-scale evidence about production behavior under user controls, strategy text, and capital.
Method
The paper analyzes two production systems sharing one design lineage, combining telemetry, registered analyses, day-clustered inference, ablations or permutation checks, and explicit evidence classes.
Results
The fleets show no directional edge: DXAP loses money and trails its matched retail benchmark, while operating-layer configuration explains major behavioral differences.
Takeaways & Limitations
The record points builders toward mechanical interventions at the order path and tool surface, including brackets, volatility-scaled leverage, and render-surface experiments.
Takeaways & Limitations
Most DXAP fills are paper trades with zero slippage and funding and a 0.4% maintenance-margin placeholder that liquidates roughly 2× later than real Hyperliquid margining.
Abstract
from arXiv · showhide
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.
1 Introduction
This paper builds a population-scale production record of autonomous trading agents across two systems sharing one design lineage, focusing on how operating-layer choices shape behavior. It explicitly separates evidence classes and reports no directional-performance claim.
- Scope: 3,505 user-funded vaults in DX Terminal Pro and user-created DXAP agents provide a continuous production record across two systems.DX Terminal Pro ran for 21 days with 7.5M invocations; DXAP extended the record to an open-ended alpha platform.
- Operating layer: The paper treats sliders, prompt templates, rendered market surfaces, and order-path mechanics as the production variables that determine agent behavior.It marks headline claims as FIRM or PROVISIONAL and records earlier retractions explicitly.
- Evidence stance: The study reports no directional-edge claim and states that the DXAP fleet loses money and trails its retail benchmark.This negative result is presented as part of the record’s scientific value rather than as a performance claim.
- Common design lineage: The shared architecture exposes five 1–5 sliders, free-text strategy, one-action-per-turn tool calls, and continuously versioned Go prompt templates.The systems share design lineage but differ in deployment mechanics: bounded real-capital tournament versus open-ended leveraged perpetuals.
- Measurement discipline: The two systems are joined at the metric level rather than row level because their design lineage does not imply schema compatibility.DXAP paper and live fills remain flagged, while cross-era P&L is restated at a common fee rate.
3 Population and Behavior Overview
The population overview places two production deployments on one measurement timeline and compares DXAP with matched Hyperliquid retail traders. The DXAP fleet traded frequently and directionally, but remained unprofitable and underperformed the benchmark.
- Measurement timeline: The record spans a February–March DX Terminal Pro window and a June–August DXAP window within one design lineage and measurement discipline.Figure 1 shows the DX Terminal Pro period as shaded and DXAP activity through daily active agents and cumulative finalized turns.
- DX Terminal Pro: 7.5M invocations and approximately 300K onchain actions came from 3,505 funded vaults during the 21-day DX Terminal Pro tournament.Behavior was bursty and social, including a one-hour FEET buying concentration involving 1,544 of 3,454 active vaults.
- DXAP versus retail: 41% of DXAP roundtrips were winners versus 50% for matched retail, while only 15% of active agents were net-positive versus 53% of retail accounts.Median hold times were nearly identical: 2.05 hours for agents and 2.18 hours for retail.
- DXAP versus retail: DXAP agents took approximately 78% long entries versus retail’s approximately 60% and used roughly 4–5× effective leverage.The chosen-leverage median was 5.0× for agents.
- Profitability: Cumulative realized DXAP P&L was −$217K at the common 5.5 bps fee rate, with no user net-positive at the July 12 census.Positions held at least 24 hours had +0.23% median account return, versus −0.15% for positions held under one hour.
4 The Operating Layer Determines Behavior
The operating layer—configuration surfaces, rendered candidate lists, and order-path mechanics—sets trading behavior more strongly than strategy text. Risk controls determine leverage, rendering choices route selection, and conflict resolution between sliders and text varies across systems.
- The operating layer sets behavior more strongly than model-written strategy text.Configuration surfaces, rendered market information, and order-path mechanics are the principal behavioral determinants in the record.
- Sliders set leverage; identity explains the rest: +0.425× leverage per riskTolerance level; agent fixed effects absorb 60% of variance.When users name leverage in strategy text, realized leverage tracks it at Spearman ρ = +0.836, while market state, recent P&L, and prompt family move sizing less.
- A render boundary causally routes selection: 1.75× selection at the rank-3/rank-4 render boundary, with a registered RDD interval of [1.49, 2.06].The discontinuity reflects a rendering choice: symbols just below the boundary have statistically identical market states but are picked less because they are not shown.
- A render boundary causally routes selection: 46.5% of entries use rendered symbols versus an 8.9% random-availability baseline, a 5.2× over-selection across postures and prompt templates.The movers surface renders the top three gainers, losers, and volume leaders.
- Strategy text vs. constraints: a cross-system contrast: Text–control conflicts resolve oppositely across systems: sliders constrain strategy text in DX Terminal Pro, while text overrides the frequency slider in DXAP.This contrast makes the winning control surface an operating-layer design choice rather than a fixed property of the model.
5 Risk Behavior
Risk behavior is governed primarily by configuration and operating surfaces rather than adaptive sizing or prompt-side controls. Volatility-blind leverage and concentrated liquidation risk coexist with a qualified, post-order-only risk signal.
- 5.1 Sizing is volatility-blind: 5.0× median leverage persists across every volatility sextile despite a 5.7× volatility spread, while realized returns worsen from −10.6 to −98.2 bps.Volatility and leverage are uncorrelated (ρ = −0.001), and notional increases with volatility.
- 5.2 Liquidation risk is concentrated in one posture-slider cell: 128 of 205 liquidations (62%) occur in the momentum × frequency-5 cell, which represents roughly 11% of the book.The cell’s day-stratified odds ratio is 22.37 [12.59, 37.45].
- 5.3 Prompt-side control fails: Requiring agents to state liquidation distance fails as a restraint: staters liquidate at 5.8% versus 1.2% for non-staters.Agents state liquidation distance in 45.0% of entry turns and size identically anyway.
- 5.4 A preregistered probe of implicit risk representations (H2): +0.0150 PR-AUC is the narrow gain from a neural encoder over 25 concrete features, but the result is qualified as PROVISIONAL.The gain survives several checks, while the TF-IDF arm scores −0.0056 at chance.
- 5.4 A preregistered probe of implicit risk representations (H2): The representation adds risk information only after the order line is read, with pre-order anchors adding nothing and post-order decoding echoing order text.The consecutive pre-order→post-order delta is +0.0179 [−0.0026,+0.0430], which does not clear zero.
- 5.4 A preregistered probe of implicit risk representations (H2): At a 10% flagging budget, the representation catches 71/205 liquidations versus 57/205 for concrete features, making it a post-order screening tool.Precision is 11.1% versus 8.9%; the paper reports no pre-decision risk signal to harvest.
6 Capture and Exits
Agents frequently reach profitable price excursions but rarely retain them, whereas explicit bracket mechanics improve realized outcomes. The measured benefit comes mainly from truncating losses rather than restricting exposure.
- 6.1 The capture gap: 43.2% of positions reached at least +300 bps favorable excursion within 24h, yet 49.3% of those closed with negative returns.Bookwide median favorable excursion was +247.9 bps against median realized −33.3 bps.
- 6.1 The capture gap: The top predicted-upside quintile reached +410 bps median MFE but realized −71.0 bps median return, so selecting predicted upside worsened P&L.Widening stops in the wildest quartile was also harmful at −68.9 [−115.5,−17.0] bps.
- 6.2 Mechanical exits beat discretionary exits: +39.0 bps per position is the paired gain from attaching a fixed 2%/4% stop/target bracket at entry.The gain remains +16.6 [+2.1,+30.8] bps after excluding liquidations.
- 6.2 Mechanical exits beat discretionary exits: Roughly 60% of the bracket gain comes from preventing blow-ups, while the remainder reflects reduced giveback on ordinary positions.The bracket preserves aggressive exposure rather than restricting it.
- 6.2 Mechanical exits beat discretionary exits: None of 16 alternative exit policies is profitable outright, with the best at −9.5 [−25.1,+8.2] bps.The paper locates exit discipline in the tool surface rather than in agent journals or stated plans.
7 Null Results
Across the measured fleets, no information source, cohort, or strategy group produces a durable directional edge. The null record also exposes chase-state over-selection, memory contamination, and a qualified restraint signal.
- Directional nulls: No directional edge appears from signals, contextual information, research recommendations, or agent selection.A 77K-candidate ledger sorts at chance, probes add only +0.001–0.003, the world-context arm goes 0/84, and selection correlation is +0.018.
- Directional nulls: No group’s lower day-clustered confidence bound exceeds zero across 8 strategy postures, 6 prompt templates, and 9 cohorts.The nearest estimate is +22.6 bps with a confidence interval of [−0.2,+47.0], so it does not clear zero.
- Behavioral nulls: 2.46× availability concentrates entries in > +0.75%/1h chase states, but forward-4h return is −6.9 bps and within-chase picks underperform random by −7.82 bps at 1h.The same negative result crosses systems in 4,096 DX Terminal Pro entries at −3.774 bps per percentage point of prior-4h side-aligned return.
- Behavioral nulls: Memory-write frequency correlates negatively with P&L at ρ = −0.200, while deleted strategies and fabricated funding income show memory contamination in production.66% of 77,269 memory operations are compactions, and 40% of compactions grow the file.
- Qualified result: A qualitative-restraint arm scores +0.52 points, but its negative-form template result remains PROVISIONAL pending replication on a fresh window.Intervention rate correlates with score at −0.72, and the authors flag template league effects as especially vulnerable to prior retractions.
8 Toward an Edge: A Development Program
The development program shifts attention from prompting and model shopping toward operating-layer interventions, utility-tested tools, environment-specific training, and recurring benchmarks. Its ordering follows the record’s measured effect sizes and evidence classes.
- Operating-layer interventions: Every measured positive lever operates on machinery around the model, including a +39.0 bps bracket and a 1.75× render-boundary selection effect.These findings motivate interventions to the order path and market rendering rather than relying on better prompts.
- Operating-layer interventions: The first pillar is atomic protection, volatility-scaled leverage at the tool, and deliberate A/B testing of the render.A mechanical trigger-quota error left 24 of 35 successful opens unprotected in one 48-hour window.
- Utility-tested tools: A tool surface remains in the manifest only if it demonstrably changes behavior, including subagents treated as tool surfaces.Structured exits pass this bar, while market-research recommendations are null at n = 120.
- Risk alignment: At a constant 10% flagging budget, post-order representation checks catch 71 of 205 liquidations versus 57 of 205 for concrete features.The paper describes this as a +25% liquidation catch and pairs it with a decision-provenance ledger.
- Model comparison: Three frontier models span 263.27–264.37 bps of replay regret, with overlapping intervals and no significant leader comparisons.Holm-adjusted tests against the leader have minimum p = 0.46, indicating no measurable decision-quality separation at this horizon.
- Model comparison: Replay bounds model differences under frozen contexts but says nothing about live P&L.The authors explicitly distinguish replay behavior from profitability claims.
- Environment-specific training: The proposed training pillar uses 231,638 finalized production turns and reports an early DX-SFT-0.3 improvement from 119 to 96 on an internal evaluation.The SFT result is early, internally reported, and explicitly not a claim of trading edge; live-branch GRPO is proposed next.
- Benchmark infrastructure: Recurring model leagues, versioned DXTradeBench signals, and public datasets are intended to keep development claims auditable.The program treats benchmarking as infrastructure because three prior claims were retracted and one rolling p-value changed from 0.0067 to 0.19 to 0.0277.
9 A Methodology Canon for Agent-Trading Research
The methodology canon converts the study’s retractions and failed claims into rules for inference, validation, measurement, and reproducibility. It emphasizes correct inferential units, preregistration, nulls, timing, and mechanism-level scrutiny.
- Purpose: Every rule in the canon was motivated by a retraction or failed claim and is presented as an enforced, transferable practice.The authors frame the canon as the most transferable output of the production record.
- Inference: The market day, never the trade, is the inferential unit because same-day positions share the tape.Position-level intervals are estimated to be approximately 2.5× too narrow.
- Inference: Pre-register the band or report the entire sweep, because one of 35 nested definitions clearing zero can arise by chance.The rule directly targets selective reporting across nested definitions.
- Validation: Require an increment over a concrete baseline; this eliminated three apparent edges that merely re-encoded visible features.The canon treats baseline comparisons as necessary evidence of added predictive content.
- Validation: Use leave-window-out cross-validation and permutation nulls rather than random splits or uncalibrated comparisons.Random folds leak across time-series data, while permutation nulls are required on every arm.
- Model evaluation: The model league compares frontier-model regret and forced-arm choice stability, with frontier models within noise at this horizon.Figure 8 uses regret against an achievable envelope and separately reports choice stability in the forced-action arm.
- Measurement: Conditioning on trades taken is descriptive only because removing the agent expands the state panel from 916 positions to 104,780 market-hours.The rule identifies post-treatment conditioning as a source of selection distortion.
- Measurement: Prefer revealed preference to outcomes: over-selection ratios omit market-variance terms and can be read in a day, unlike outcome metrics requiring weeks.The canon prioritizes faster behavioral measurements when they directly capture selection.
10 Related Work
Related work spans live-capital LLM trading tournaments, finer-grained trading evaluations, the authors’ earlier production architecture, and benchmark or harness studies. This paper’s distinction is a population-scale production record with operating-layer telemetry.
- Agentic trading evaluations: Public LLM-trading leaderboards pit frontier models against live capital, but function as tournaments that publish P&L without operating-layer telemetry.The paper contrasts this tournament format with its own measurement-program framing.
- Agentic trading evaluations: KTD-Fin and Agent Market Arena evaluate trading behavior with finer task decomposition or arena-style comparison, extending the evaluation landscape beyond simple leaderboard P&L.The passage places these works within the broader agentic-trading evaluation literature.
- Prior work: Relative to the authors’ prior work, this paper adds trace-level analysis, the DXAP fleet record, cross-system findings, and production failure measurements.The prior work introduced DX Terminal Pro’s architecture and failure modes while deferring deeper user-to-agent-to-execution trace analysis.
- Benchmarks and harnesses: MEMEbench and harness-transfer work study model selection and EVM construction accuracy, whereas this paper takes the harness as given and measures the population under it.The distinction is between benchmark or harness capability studies and production population measurement.
11 Limitations
The study’s cross-era comparisons are bounded by paper-engine mechanics, thin liquidation-tail evidence, single-venue windows, concentrated model mixes, and restricted raw telemetry.
- Paper-engine scope: 5,035 DXAP fills used live capital, but most fills had zero slippage, zero funding, and a 0.4% maintenance-margin placeholder.Liquidation-related results should therefore be interpreted as earlier and larger under real margining.
- Statistical scope: 205 liquidations support the tail results; the concentration finding survives stratification, but cell counts remain small.
- External validity: The evidence spans Base memecoins in one window and Hyperliquid perpetuals in another, so cross-venue generality is asserted only for findings crossing both systems.
- Model mix: DX Terminal Pro uses one frozen model, while 86 of 91 active DXAP agents were qwen3.7-plus on July 22.
12 Conclusion
Across two production systems, agent behavior was shaped more by the operating layer than by model reasoning, while the fleets showed no directional edge. The paper therefore points toward mechanical operating-layer interventions and a methodology canon grounded in measured failures and retractions.
- Core conclusion: The operating layer shaped behavior more than model reasoning: leverage stayed fixed as volatility varied 5.7×, a render boundary moved selection 1.75×, and one posture cell held 62% of liquidations.
- Core conclusion: Neither production fleet produced a directional edge, while the record identified volatility-blind sizing, render-driven chase, and uncaptured excursion.
- Implications: The paper recommends mechanical interventions: bracket at the order path, scale leverage to realized volatility at the tool, and A/B test the render.
- Evidence: Aggregate statistics supporting every headline claim appear in the text and figures, with public datasets, dashboards, and per-claim evidence artifacts published online.