Source-linked AI summary

RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents

Yupeng Zhang, Liuyuan Jiang, Hongyi Huang, Bingheng Li, Lisha Chen

arXiv:2608.28399v1cs.AIq-fin.TR

TL;DR

Sequential LLM trading policies may become predictable when they respond systematically to price movements, but their directional structure is not captured by aggregate task scores alone. This paper introduces RetailAgent to evaluate anonymized intraday long/flat decisions against subsequently revealed returns under a controlled information boundary. The agents exhibit consistently negative exposure-matched timing across tested configurations; shuffling attenuates the effect, while self-authored memory increases persistence and conditional adverse timing.

  • Problem

    The paper asks whether sequential LLM trading decisions exhibit stable directional structure that could be anticipated by other market participants.

  • Method

    RetailAgent gives an LLM anonymized intraday price histories and permitted endogenous state, then evaluates repeated long/flat choices against subsequent research-return labels.

  • Results

    Within-stock timing is consistently negative across 14 configurations spanning modality, decision horizon, account state, and model family.

  • Takeaways & Limitations

    The findings reveal stable, recoverable directional structure in sequential LLM action traces beyond average exposure and tested price-based signals.

  • Takeaways & Limitations

    The evidence concerns controlled LLM policies on anonymized historical equity paths with fixed prices and research-return labels; human traces and executable-return analysis remain separate requirements.

Abstract

from arXiv · show

In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants. This paper studies whether large language model (LLM) agents exhibit such directional structure through RetailAgent, an experimental framework in which an LLM observes anonymized intraday equity price histories and permitted state, then repeatedly chooses long (hold the stock) or flat (stay out) before the subsequent interval return is revealed. We compare returns during long and flat intervals along the same stock's intraday path after removing the overall fraction of long decisions. This exposure-matched measure reveals persistent negative timing across modality, horizon, state, and model family. Shuffling saved action sequences substantially attenuates the effect, showing that alignment between actions and subsequent returns drives the negative score. Feeding self-authored memories into decisions further increases policy persistence, while timing becomes more negative among stock-days on which the agent uses both actions. These results reveal stable, recoverable directional structure in sequential LLM financial decisions and a behavioral signal for studying how another participant could respond to a predictable policy.

1 Introduction

RetailAgent is introduced as a controlled multimodal framework for testing whether sequential LLM trading decisions contain stable directional structure. Across configurations, timing is consistently negative, trajectory shuffling attenuates the effect, and self-authored memory increases persistence and adverse timing.

  • Framework: RetailAgent traces sequential LLM decisions and self-authored memory under a fixed information boundary using an anonymized intraday price interface.The agent receives permitted price history and endogenous state, chooses long or flat, and records subsequent research returns.
  • Adverse timing: Across 14 standard configurations spanning modality, horizon, account state, and model family, within-stock timing is consistently negative.In principal 10-minute Qwen3.5 panels, numerical-text, chart-only, and joint text+chart inputs score −45.7, −29.9, and −48.9 bps per stock-day, respectively, on different sampling frames.
  • Validation: Trajectory shuffling reduces the intact-schedule estimate of −45.7 bps to −3.5 bps under global shuffling and −8.7 under same-day shuffling.Preserving the overall long fraction while permuting actions shows that sequential alignment with subsequent returns accounts for most of the intact estimate.
  • Self-conditioning: Self-conditioning reduces action switching and makes within-stock timing more negative among stock-days containing both actions.With one prior self-authored memory visible, timing shifts from −62.8 to −74.1 bps and position changes from 2.5 to 1.7 per stock-day.
  • Positioning: The framework complements research on financial LLMs and sequential trading agents by auditing a fixed-path long/flat policy.Its architecture combines text, vision, self-authored memory, and an optional learned branch that supplies continuous embeddings to a frozen LLM.

2 Methodology

RetailAgent evaluates sequential long/flat decisions on anonymized intraday equity paths under a fixed information boundary. Its exposure-matched timing metric separates passive exposure from whether the agent holds during relatively favorable intervals, while controlled conditioning varies representations and endogenous state.

  • Controlled decision loop: RetailAgent records each observation, endogenous state, binary action, self-authored memory, and subsequently revealed return along anonymized intraday stock paths.The protocol conceals stock identity, date, absolute price, news, fundamentals, and cross-sectional ranks.
  • Timing metric: Exposure-matched timing measures whether the agent is long during relatively favorable intervals on the same stock-day after removing average long exposure.A policy that is always long scores zero by construction; negative values indicate wrong-signed timing.
  • Timing metric: Timing alpha decomposes long-only cumulative return into passive exposure to the return path plus a centered timing term.This permits positive long-only return to coexist with systematically adverse within-stock timing.
  • Controlled interventions: The framework holds binary long/flat semantics and the market information set fixed while varying observation construction and carried endogenous state.Conditions include numerical text, charts, multimodal inputs, account state, narrative memories, and an optional latent price-dependent channel.
  • Evaluation: The complementary schedule 1 − p reverses timing alpha by definition, providing a sign diagnostic for the timing measure.Timing evaluates when a stock is held, whereas contemporaneous IC evaluates which stocks are preferred at the same decision time.
  • Controlled interventions: Narrative self-conditioning exposes the decision to a current memory and up to w earlier self-authored memories before choosing long or flat.The evaluated memory windows are w ∈ {0, 1, 5, 12}; the optional latent channel uses a GRU price encoder and soft-token embeddings.

3 Experiments

RetailAgent evaluates sequential long/flat LLM decisions on anonymized historical equity paths and finds consistently adverse within-stock timing across experimental configurations. Controls and self-conditioning analyses show that sequential alignment and persistent self-authored memory are central to the observed pattern, while the evidence remains bounded to fixed historical paths and research-return labels.

  • Experimental setup: RetailAgent evaluates anonymized intraday price histories, permitted account state, and self-authored memory across modality, horizon, state, and model-family conditions.The framework records observations, endogenous state, actions, memories, and subsequently revealed returns under a controlled price-only interface.
  • Standard timing: Across 14 standard configurations, within-stock timing remains negative, indicating unfavorable entry and exit intervals relative to each stock’s own path.The principal 10-minute Qwen conditions score −45.7 bps for text, −29.9 for chart, and −48.9 for multimodal inputs, using distinct sampling frames.
  • Controls: Shuffling actions reduces timing from −45.7 bps intact to −3.5 bps globally and −8.7 bps within day, isolating sequential action-return alignment.The complementary schedule reverses the timing sign by exchanging long and flat actions and serves as a sign diagnostic.
  • Standard timing: −46.2 bps is the difference between the LLM schedule and exposure-matched random schedules on the principal text sample, with 95% CI [−50.5, −41.8].Exposure-matched random schedules provide the behavioral benchmark while price-based signals serve as overlap comparators.
  • Signal overlap: Residual timing remains positive after linear projection on conventional price signals, retaining 43.9 of 46.8 bps for GRU, 40.4 of 45.7 bps for GBDT-38, and 38.9 of 45.7 bps for reversal.Correlations with the inverted schedule remain below +0.3, so the residuals quantify structure unexplained by each projection under research labels.
  • Self-conditioning: One earlier self-authored memory shifts timing from −62.8 to −74.1 bps and reduces position changes from 2.5 to 1.7 per stock-day among stock-days containing both actions.Across the self-conditioning sweep, main-period timing intervals remain below zero, while timing and contemporaneous cross-sectional IC represent separate empirical axes.

4 Conclusion

RetailAgent reveals consistently negative exposure-matched within-stock timing across tested configurations. Shuffling attenuates the effect, while self-conditioning increases persistence and makes conditional timing more negative.

  • Negative exposure-matched within-stock timing persists across input modality, decision horizon, account state, and model family.
  • Trajectory shuffling substantially attenuates the timing effect, indicating that sequential alignment between actions and subsequent returns drives the intact estimate.
  • Self-conditioning increases policy persistence and shifts conditional timing among stock-days containing both long and flat actions.
  • The findings identify stable directional structure in LLM action traces beyond average exposure and the tested price-based signals.

Supplementary Material

The appendices document the protocol, supporting analyses, and provenance underlying the paper’s main numerical comparisons.

  • Appendix A records the experimental protocol and reproducibility information.
  • Appendix B presents visual summaries and supporting analyses.
  • The main paper retains principal numerical comparisons, while the appendices document how each estimate was produced.

A Expanded experimental protocol

The expanded protocol defines anonymized intraday stock-day inputs, long-versus-flat decisions across multiple cadences, and stateful or multimodal conditions. Prompts specify the required decision format and condition-specific information supplied to the agent.

  • The source data represent each anonymized stock-day on a 239-minute intraday grid with indexed serialization preserving relative price movements.
  • The primary local panel scores 23 10-minute intervals per stock-day, while 5-minute and 20-minute panels score 47 and 11 intervals.
  • Agents decide whether to be long or flat and must return a memo followed by an exact binary ACTION line.
  • Stateful conditions add position information, trade history, or recent self-authored memories to the task and response format.
  • Chart-only and joint conditions provide structured multimodal messages with an image block preceding the text block.
  • The joint condition combines a rolling one-minute chart with the numerical list used by the text condition.

A.3 Generation and parsing

Generation uses a locally cached checkpoint with specified sampling and parsing infrastructure, while the archived implementation contains horizon-wording inconsistencies and documents failed-interval handling. The price encoder and adapter are defined with explicit architectural and pretraining settings.

  • Generation and parsing: The principal conditions use the locally cached Qwen/Qwen3.5-9B checkpoint through vLLM with thinking disabled.
  • Generation and parsing: Numerical and state conditions use a 160-token output limit, whereas chart and joint text-chart conditions use a 320-token limit.
  • Generation and parsing: Archived horizon wording is inconsistent: 20-minute state conditions display 10-minute wording, while the visual generator produces 21-minute wording on the 11-bar panel.
  • Generation and parsing: The price encoder is a two-layer GRU with 256 hidden units, a 256-dimensional output, dropout 0.1, and a 120-minute lookback.
  • Generation and parsing: The narrative provenance table flags cells containing failed intervals and reports scored stock-days used for archived conditional timing estimates.
  • Generation and parsing: VICReg pretraining uses eight epochs, batch size 512, and 60 stocks selected with NumPy seed 0.

A.5 Inference and metric provenance

Timing alpha is computed by aggregating within stock-days and averaging across eligible trajectories, while Figure B1 visualizes the corresponding estimates and uncertainty.

  • Timing alpha aggregates interval-level values within each stock-day before averaging across stock-days admitted by the scorer.
  • All-long and all-flat trajectories receive zero timing, so reported means characterize the switching subset.
  • Figure B1 reports within-stock timing alpha with 95% date-clustered confidence intervals for standard and representative self-conditioning conditions.

A.6 Reproducibility record

The accompanying artifact provides code, records, checkpoints, counters, and machine-readable outputs intended to reproduce the reported evaluations.

  • The artifact includes panel-construction and scoring code, prompt implementations, aggregate trajectories, hosted-model records, checkpoints, failure counters, and evaluation outputs.
  • Saved positions reproduce reported scores, while generation code and aligned panel values reconstruct the prompts.
  • Narrative trajectory files preserve parsed actions and one truncated midpoint memory excerpt per stock-day.

B Additional experimental and analysis details

Figures B1 and B2 complement the numerical tables by making sign recurrence, uncertainty, and the distinction between timing and cross-sectional selection easier to inspect.

  • Figures B1 and B2 provide visual counterparts to Tables 1 and 3, while the tables remain the numerical record.
  • Figure B2 reports within-stock timing alpha and contemporaneous cross-sectional IC across ten Qwen3.5-9B narrative self-conditioning conditions.
  • Each Figure B2 condition attempts 1,500 stock-days, with scored counts reported in Table 3 and 95% date-clustered confidence intervals in both panels.

B.1 Latent-state and metric analyses

The analyses separate configurations by objective comparability and quantify how timing, IC, and portfolio return relate across experimental conditions.

  • B.1 Latent-state and metric analyses: Table 5 groups the disabled-token, two-token, eight-token, sixteen-token, and 492-word variants under one cross-sectionally centered outcome.
  • B.1 Latent-state and metric analyses: Trainable-encoder, frozen-encoder, free-vector, combined-transformation, and memo-conditioned rows use separate transformed objectives or evaluation frames.
  • B.1 Latent-state and metric analyses: Timing and IC correlate 0.684 across 12 conventional conditions and −0.075 across ten narrative conditions.
  • B.1 Latent-state and metric analyses: Portfolio return and IC correlate 0.859 across conventional conditions and 0.920 across narrative conditions.

B.2 Archived trajectory and cross-model records

Archived trajectory examples illustrate how action persistence can coincide with adverse within-path timing, while cross-model records extend the negative-sign pattern beyond the principal experiments. The paper cautions that these findings are controlled historical-path labels rather than economic trading evaluations.

  • Archived trajectory: Seven long intervals summed to −576.7 bps, while sixteen flat intervals summed to +575.9 bps, producing timing alpha of −576.4 bps.This price-text illustration is an ex-post example of the paper’s timing measure.
  • Archived trajectory: A consistency example preserved a bullish memory across decisions 8–23, with exposure 0.696 and timing alpha of −867.3 bps.The trajectory contained one position change; long-interval return was +359.4 bps and flat-interval return was +1,404.0 bps.
  • Cross-model records: Claude Haiku 4.5 produced negative estimates at both tested horizons, using 827 stock-days at 5 minutes and 165 at 10 minutes.These are cross-model sign checks reported in Table 1.
  • Cross-model records: Granite 4.0 produced negative point estimates across price-only, position, ledger, and account prompts on 600 stock-days per condition.The estimates were recovered from per-cell records after a cache migration.
  • Validity conditions: The outcomes are research-return labels from anonymized historical paths, not economic evaluations incorporating execution, borrowing, latency, capacity, or market impact.The paper identifies reserved-market evaluation and interactive-market extensions as relevant additions.
Loading 2608.28399v1…