Source-linked AI summary

Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing

Linsen Zhu, Mengqing Cai

arXiv:2609.04917v1cs.AIq-fin.PMq-fin.TR

TL;DR

This review asks whether AI’s technical progress in investment workflows produces sustainable net alpha and evaluates that question across equities, ETFs, and several crypto market settings. Using an alpha-translation chain, it finds substantial upstream progress but no general architecture shown to deliver persistent, cross-regime, capacity-aware net alpha, while identifying conditions that could make future claims more credible without guaranteeing profit.

  • Problem

    AI studies document technical progress, but do not by themselves establish whether that progress produces sustainable net alpha.

  • Method

    The review evaluates public research across investment workflows using an alpha-translation chain from point-in-time information through signals, positions, execution, and risk-adjusted returns after costs.

  • Results

    The examined record shows material technical progress upstream, but does not establish a general AI architecture producing persistent, cross-regime, capacity-aware net alpha.

  • Takeaways & Limitations

    More credible claims require point-in-time data and models, decision-aligned objectives, joint portfolio–execution evaluation, controlled adaptation, prospective records, and authority-matched governance, none of which guarantees profit.

  • Takeaways & Limitations

    The evidence is constrained by implementation and market-mechanics issues, including capacity, simulation assumptions, venue rules, execution, and feedback from other investors’ adoption.

Abstract

from arXiv · show

Artificial intelligence (AI) now supports investment workflows from data and prediction through research, portfolios, execution, and tool use. Technical capability, however, is not evidence of investment profitability. This critical state-of-the-art review examines public research available through 31 August 2026 on listed equities, exchange-traded funds, centralized crypto spot, perpetual futures, and on-chain markets. We organize evidence with an alpha-translation chain: point-in-time information must yield a stable signal, feasible positions, executable orders, and risk-adjusted returns after costs. Across machine learning, time-series foundation models, financial language models, reinforcement learning, and agents, the examined record shows real but mainly upstream progress in prediction, text processing, portfolio design, and workflow integration. Evidence is thinner for durable net performance. Temporal contamination, repeated selection, survivorship, weak benchmarks, implementation costs, venue mechanics, and capacity can break translation to net alpha. Strong historical results coexist with predictor decay, corrected look-ahead failures, mixed prospective evidence, and few audited live-capital records. Crypto adds informative state but requires separate treatment of spot, perpetual, and decentralized cash flows and execution. Within the public evidence examined here, no general AI architecture is shown to deliver persistent, cross-regime, capacity-aware net alpha. More credible claims require point-in-time data and models, decision-aligned objectives, joint portfolio--execution evaluation, controlled adaptation, prospective tests, and authority-matched governance. These conditions can improve evidence and implementation; they do not guarantee profit.

1 Introduction

AI now spans the investment workflow from information processing and prediction to portfolio construction, execution, and agents, but technical progress does not establish sustainable net alpha. This review therefore evaluates whether information survives translation into feasible, executable, risk-adjusted returns across distinct market structures.

  • AI across investment workflows: AI systems increasingly connect prices, text, blockchain state, forecasts, holdings, and order-routing tools within one investment workflow.The literature includes nonlinear return models, cost-aware portfolios, language-based signals, LLMs, and agentic tool use.
  • From prediction to profitability: Prediction accuracy or positive backtest returns do not by themselves establish alpha, net performance, or prospectively attainable trading results.Capacity, market impact, borrowing, execution delay, selection multiplicity, and adoption can weaken apparent performance.
  • Alpha-translation chain: The review evaluates AI through an alpha-translation chain linking point-in-time information, stable signals, constrained positions, plausible fills, and post-cost returns.The chain also requires persistence across regimes, capital scales, and independent or prospective tests.
  • Market-specific evidence: Equities, ETFs, centralized crypto spot, perpetual futures, and on-chain markets require distinct evaluation because their trading, settlement, financing, and execution mechanics differ.Crypto introduces continuous venues, funding, liquidation, custody, smart-contract, and blockchain-state considerations.
  • Bounded conclusion: The examined public record shows meaningful technical advances but does not establish a general AI method with persistent, cross-regime, capacity-aware net alpha.Historical successes coexist with fragile signals, concentrated portfolios, and prospective results without statistically significant abnormal returns.
  • Review contribution: The review synthesizes established financial disciplines across AI methods and market strata rather than treating any single task metric or backtest as decisive.Its evidence profile distinguishes prediction, simulation, prospective evaluation, live-capital operation, and external persistence.

2 Scope, Review Method, and Evidentiary Boundaries

This critical state-of-the-art review defines a broad but bounded AI-and-asset scope and evaluates claims through a multidimensional evidence framework. It is intentionally selective rather than systematic, and it treats missing documentation as unknown rather than favorable.

  • Review method: The paper is a critical state-of-the-art review using iterative searches, citation tracing, publisher records, and official documentation rather than a database-complete systematic review.Primary studies and official proceedings are prioritized, while vendor documentation establishes capabilities rather than return performance.
  • Method limitations: The review does not claim exhaustive retrieval, duplicate screening, a registered protocol, formal risk-of-bias scoring, or valid pooled effect estimates.Its synthesis retains contrary findings and separates inference from directly reported results.
  • Scope: The review covers AI methods that bear on investment research, allocation, trading, execution, or agents, while excluding unrelated financial applications.AI is an umbrella term; prompted workflows are not labeled RL without sequential learning, and tool use is not autonomy without delegated authority.
  • Scope: Its asset boundary includes listed stocks and ETFs, centralized crypto spot, perpetual futures, and on-chain or DEX markets, but excludes most other asset classes.ETFs are included as investable portfolios and execution vehicles for asset-allocation policies.
  • Evidentiary framework: Its evidence framework records temporality, selection, portfolio mapping, implementation, risk and benchmark, external validity, and operational provenance.These dimensions distinguish retrospective simulations, prospective paper portfolios, and documented live capital.
  • Evidentiary boundaries: Reported omissions are coded as unknown rather than favorable, so claims such as “net” remain limited to disclosed cost models and assumptions.The review also rejects inferring zero latency, unlimited borrowing, safe operation, or investment skill from silence or connectivity alone.
  • Interpretive boundaries: Multiple testing and heterogeneous trading costs prevent universal conclusions from either failed replications or optimistic cost estimates.The review therefore requires higher discovery thresholds and sensitivity to investor size, participation, urgency, and scale.
  • Core definitions: Sustainable net alpha requires risk-adjusted performance after explicitly modeled costs to survive time-valid external or prospective evaluation across regimes at a stated scale.This standard is narrower than claims that any AI strategy has never earned money or that proprietary systems cannot work.

3 From Model Output to Sustainable Net Alpha

The paper treats profitability as a conversion process: point-in-time information must become a stable signal, feasible holdings, executable orders, and persistent net returns. Evidence quality therefore depends on multiple dimensions rather than model output or backtest performance alone.

  • Alpha-translation chain: A representation from point-in-time information can support forecasts, rankings, event labels, or distributions, but it becomes economically meaningful only after portfolio and execution decisions.The framework separates information, representation, policy, target holdings, orders, fills, and realized holdings.
  • Alpha-translation chain: Net performance deducts financing and venue-specific costs, including borrow, leverage, perpetual funding, liquidation, gas, failed transactions, MEV, fees, spreads, and impact.Costs are return-normalized for the same measurement period and may depend on orders, fills, and holdings.
  • Alpha-translation chain: Sustainable net alpha requires persistence outside the researcher-selected environment and at a stated capital scale.A profitable simulation can still fail when signals decay, borrow is unavailable, fills arrive late, or adoption reduces opportunity.
  • Multidimensional evidence profile: Evidence quality is a seven-dimensional profile covering temporality, selection control, portfolio mapping, implementation realism, risk and benchmarks, external validity, and operational provenance.The dimensions are not interchangeable, and evaluation-design archetypes are not ranked levels.
  • Multidimensional evidence profile: An out-of-sample label does not ensure point-in-time validity because later revisions, future survivors, or pretrained memories can contaminate inputs.The paper requires information used at the decision to exclude information observable only afterward.
  • Multidimensional evidence profile: Cost deductions are models rather than proof of realism, and cost-aware portfolio learning can change which information is valuable.Cost-unaware predictors may favor fleeting, low-scale opportunities and unattractive implementable frontiers.
  • Central verdict: The reviewed evidence shows upstream progress while leaving downstream claims about timestamped commitment, real execution, complete costs, risk adjustment, capacity, and persistence unresolved.This mismatch supports material changes in investment practices without establishing a general profit engine.

4 Prediction and Representation: Real Progress, Conditional Economic Value

Prediction and representation methods show real progress in extracting nonlinear, high-dimensional information, but their economic value remains conditional on temporal validity, implementation, and decision alignment. Foundation models improve forecasting infrastructure, yet direct profitability evidence remains limited.

  • Cross-sectional prediction: Nonlinearities and high-dimensional interactions contain information under established equity-model protocols, but these results do not establish universal alpha.Historical universes, rebalancing rules, and institutional assumptions define the opportunity set; prospective, capacity, and live-operation evidence remains open.
  • Replication and temporal validity: Published-factor evidence is sensitive to feature definitions, accounting lags, microcap treatment, weighting, exchanges, sample periods, and statistical thresholds.A neural network trained on a disputed panel can fit the disagreement rather than resolve it.
  • Replication and temporal validity: Predictive relationships can deteriorate after sample end, publication, or regime change, consistent with overfit and trading that attenuates mispricing.These risks make out-of-sample labels insufficient without continued temporal validation.
  • Decision-aligned learning: Decision-focused policies report stronger historical certainty-equivalent returns or more implementable portfolios than forecast-first approaches in their tested settings.Cost-aware learning changes portfolio composition and the attainable frontier when fleeting signals trade poorly after costs.
  • Time-series foundation models: Time-series foundation models provide transferable forecasts, but forecasting benchmarks and task-level wins do not demonstrate tradable excess returns.In five liquid US equities, pretrained TSFMs won eight of ten tasks, yet significant gains appeared in only two model–asset comparisons; neither finance preprint supplied costs, capacity, or live returns.
  • Time-series foundation models: Large pretrained language backbones are not intrinsically necessary for numerical forecasting, and numerical pretraining can create difficult-to-audit point-in-time contamination risks.A defensible investment experiment requires checkpoints whose training data ends before the trading period or a strong contamination test.

5 Text, Multimodality, and Agents as an Integration Layer

Financial language models and multimodal agents improve the extraction, organization, and integration of heterogeneous information, but these capabilities do not by themselves establish durable trading profitability. The central evidence constraint is temporal and operational provenance.

  • Text information: Financial text can contain return-relevant information when labels, release times, and audience expectations are aligned with the economic task.Early sentiment and supervised text-signal studies remain conceptually important despite preceding generative LLMs.
  • Text information: Financial language models broaden classification, extraction, summarization, question answering, and multimodal research workflows without necessarily producing trades.Analyst productivity should be measured through time saved, coverage, error rates, and decision quality rather than recast as alpha.
  • Temporal validity: Event-time-aware LLM tests find informative headline classifications, especially for smaller firms and negative news, while predictability weakens as adoption spreads.This evidence separates semantic processing from historical recall but remains a study-specific profitability result.
  • Temporal validity: LLMs can recall pretraining-period financial events, so masking identifiers alone does not prevent historical look-ahead; point-in-time checkpoints provide a stronger design.Historical tests require model information cutoffs, not only document-level timestamping.
  • Multimodality: Multimodal systems can combine text, prices, charts, audio, and blockchain state, but each added modality introduces timestamps, revisions, missingness, and provenance risks.Comparisons should hold the information set, latency, and search budget fixed before crediting architectural integration.
  • Agents: Agents integrate retrieval, memory, debate, tools, position generation, and feedback, yet correlated agents, regime anchoring, and outcome reuse can undermine independent evidence.Benchmarking is improving through temporally structured settings and real-time commitments, but complete portfolio, execution, and risk integration remains uncommon.
  • Operational authority: Trading interfaces can expose account, order, wallet, and blockchain operations, but documented capability and adoption are not evidence of profitable or safe use.Authority-bearing systems require controls attached to permissions, while provenance and reproducibility must survive generation and tool use.

6 Portfolio Learning, Reinforcement Learning, and Risk

Portfolio learning and reinforcement learning address weaknesses of forecast-first investing by modeling constraints, dynamic costs, execution, and risk. Their promise remains bounded by objective specification, simulator fidelity, benchmark choice, and multidimensional deployment risk.

  • Portfolio learning: End-to-end portfolio policies optimize holdings under investor objectives rather than minimizing intermediate forecast error, with reported historical gains in tested settings.Direct policies can learn nonlinear feature-to-weight mappings, while risk aversion regularizes them economically.
  • Portfolio learning: Cost-aware learning treats signal decay, turnover, and implementation as portfolio-design variables rather than post hoc evaluation statistics.Learning the implementable efficient frontier can alter portfolio composition when stronger short-lived forecasts are consumed by spread and impact.
  • Reinforcement learning: Reinforcement learning is appropriate when actions change future states or opportunities, including inventory, execution costs, tax lots, leverage, liquidation risk, and information leakage.Reproducible environments and standardized datasets address inconsistent preparation and evaluation, but do not remove market-simulation assumptions.
  • Reinforcement learning: Historical RL environments can reward simulator shortcuts because they omit counterfactual prices and queues and often assume price-taking, next-bar fills, or future-fitted normalization.Repeated interaction with a fixed historical tape is not interactive market evidence.
  • Crypto evaluation: Crypto RL results vary with period, benchmark, action space, reward, and cost model, as contrasting studies report systematic frameworks versus failure to beat buy-and-hold.Interpretation also requires volatility-matched exposure, cash, momentum, carry, funding, and liquidation comparisons.
  • Risk: Portfolio risk is multidimensional, spanning exposure, concentration, liquidity, drawdown, leverage, counterparty, custody, stablecoin, smart-contract, and operational risks.A single scalar penalty can miss asymmetric losses and liquidation discontinuities; reporting should include net returns, drawdown, turnover, leverage, concentration, beta, factor alpha, and tail loss.
  • Execution: Execution policies should be evaluated with implementation shortfall, fill probability, adverse selection, latency, rejected orders, and tail outcomes, while enforcing hard completion, credit, and price controls.Average shortfall improvement is insufficient if risk-reducing orders sometimes fail to complete.
  • Synthesis: Decision-aligned objectives can address forecast-first weaknesses, but they relocate model risk into the objective, simulator, and constraint set.They improve on proxies only when the decision environment is time-valid and economically faithful.

7 Market Structure and the Meaning of Profitable Prediction

Profitability depends on market-specific data, cash flows, execution, and capacity rather than prediction alone. Equities, centralized crypto, perpetual futures, and on-chain markets therefore require distinct validity and P&L tests.

  • Listed equities and ETFs: Equity evidence must account for delistings, publication lags, corporate actions, historical constituents, and capacity concentrated in microcaps.Equal-weighted results can support little capital, while value-weighting changes the economic question.
  • Listed equities and ETFs: ETF rotation can make allocation views implementable, but benchmark discipline is essential because returns may primarily reflect market beta, duration, size, or geography.Leveraged and inverse ETFs also have path-dependent daily objectives.
  • Centralized crypto spot: Crypto spot offers structured signals and price segmentation, but changing universes, wash trading, custody, settlement, and stablecoin exposure complicate net performance.Prefunded cross-venue arbitrage additionally bears counterparty and withdrawal risk.
  • Perpetual futures: Perpetual-futures P&L includes funding, trading costs, margin, mark prices, and discontinuous liquidation losses, so last-trade backtests can describe unsurvivable account paths.Point-in-time contract specifications are necessary because funding intervals and limits can change.
  • On-chain markets: On-chain execution depends on reserves, fees, price impact, transaction ordering, gas, failed calls, and changing contract state rather than contemporaneous mid-prices.Liquidity provision must compare fees with rebalancing, hedging, gas, and adverse selection.
  • Cross-market implications: Across market strata, shared validation requires point-in-time universes, executable benchmarks, risk adjustment, and capacity, while marginal cash flows and failure events differ.A single return equation and cost scalar cannot represent every market.

8 Does the Evidence Establish Sustainable Excess Returns?

The reviewed record contains credible evidence of AI progress and economically meaningful historical results, but it does not establish general sustainable excess returns. Time contamination, search, implementation, equilibrium effects, and restrained prospective evidence limit the inference from favorable studies.

  • Affirmative evidence: Nonlinear equity models, direct portfolio methods, and point-in-time text models report improved historical prediction, portfolio outcomes, or economically meaningful signals.These findings reject both the claim that flexible models add no information and the claim that every favorable AI backtest is leakage.
  • Affirmative evidence: Agents raised explained contemporaneous earnings-announcement return variation from near 8% for a standard benchmark to close to 20%, but this is not trading return evidence.The result supports text-based mechanism structuring within that protocol.
  • Affirmative evidence: Crypto studies identify momentum, attention, factor structure, price segmentation, and futures carry, but these opportunities can compensate crash, leverage, custody, liquidation, and competition exposures.Their existence is an input to an AI strategy, not evidence of net returns.
  • Why evidence thins: Look-ahead alignment can erase reported alpha, while repeated prompts, models, assets, windows, and variants can inflate the observed maximum.Historical LLM evaluations are especially vulnerable when the base model has seen the outcome.
  • Why evidence thins: Costs, turnover, liquidity, capital, investor learning, and common adoption can erode signals and reduce their scalability after publication.Cost-agnostic learning may select precisely the least scalable opportunities.
  • Prospective and field evidence: Prospective household-portfolio evidence finds concentrated style tilts without statistically significant abnormal returns, while AI-fund outperformance declined over time among early adopters.These studies support potential and diffusion limits rather than a zero-effect verdict.
  • Bounded verdict: Evidence is strong for representation and selected historical forecasts, mixed for prospective decisions, and sparse for independently audited live-capital performance across regimes and scales.Accordingly, the record does not support a general sustainable-net-alpha claim.

9 A Research Agenda for More Credible and More Durable Net Returns

A more credible research program must connect point-in-time information, decision-aligned objectives, execution, adaptation, statistical validation, and governance. These measures can expose trade-offs and improve evidence, but they are hypotheses rather than guarantees of profit.

  • Point-in-time evidence: Every observation should carry event and availability times, with vintages, changing universes, historical specifications, on-chain state, and feature lineage preserved.Point-in-time base-model checkpoints are also needed for historical LLM evaluation.
  • Point-in-time evidence: Strict point-in-time standards may reduce headline performance and sample length while converting hidden bias into visible uncertainty.Shared, versioned testbeds across equities and crypto strata would support cumulative knowledge.
  • Decision-aligned objectives: Targets should match the action and horizon, with decision-focused objectives incorporating risk, turnover, costs, and uncertainty rather than relying on forecast error alone.Distributional forecasts and calibrated abstention are preferable when position size depends on uncertainty.
  • Decision-aligned objectives: Evaluation should expose holdings and market states, perturb risk and cost assumptions, compare simple policies, and report return, risk, turnover, concentration, capacity, and tail loss.Multi-objective reporting is safer than optimizing one post-selected Sharpe ratio.
  • Joint portfolio–execution evaluation: Portfolio and execution systems should pass cost and capacity upstream, report performance over capital, participation, and latency, and use separate engines for spot, perpetual, and on-chain P&L.Perpetual agents need funding and liquidation state, while DEX agents need ordering and failed-gas simulation.
  • Joint portfolio–execution evaluation: Execution policies should be stress-tested across response models and graduate through paper, shadow, and limited-live stages because simulators are inevitably misspecified.This addresses endogenous market impact rather than evaluating against a fixed tape only.
  • Controlled adaptation: Adaptation requires predeclared triggers, shadow comparison, approval, rollback, immutable version records, and attribution of performance changes.Alternative adaptation mechanisms should be compared under matched information and compute budgets.
  • Prospective validation: Crowding and search must enter evaluation through capacity changes, signal overlap, trade correlation, development ledgers, sequestered confirmation, and multiple-testing control.A static historical edge is not a stable resource when capital and public variants grow.

10 Conclusion

The review finds substantial technical progress across the investment chain, but the public profitability record remains narrower and does not establish general persistent net alpha. It recommends evaluating progress by the weakest link between information and realized performance, using more falsifiable and market-aware evidence standards.

  • Technical progress: AI advances span nonlinear equity modeling, economic-objective portfolio policies, representation transfer, language processing, agentic tool use, and crypto-market analysis.Crypto also introduces funding, liquidation, manipulation, custody, gas, MEV, and smart-contract risks that equity-style cost measures cannot capture.
  • Profitability evidence: Strong historical out-of-sample results coexist with anomaly decay, multiple-testing risk, corrected timestamps, memorization, cost-sensitive frontiers, and weak prospective evidence.The reviewed record is also limited by few independently audited live-capital results.
  • Profitability evidence: The evidence examined through 31 August 2026 does not establish a general AI architecture delivering persistent, cross-regime, capacity-aware net alpha.This conclusion does not imply that no individual strategy can earn excess returns.
  • Evidence standards: More credible claims require point-in-time data and model versions, decision-aligned objectives, joint portfolio–execution design, market-specific P&L, controlled adaptation, prospective records, and authority-matched governance.These practices make claims falsifiable and can improve evidence and implementation, but they do not guarantee profit.
Loading 2609.04917v1…