Source-linked AI summary

The Axiomatic Trader: Latent Regularity, Information Budgets, and the Canonical Form of a Quantitative Investment System

Jiayu Li

arXiv:2608.23416v1cs.LGq-fin.PM

TL;DR

The paper formalizes persistence in systematic trading through declared assumptions about latent states and representations, then derives a constrained architecture and tests its falsifiability. Its empirical evidence finds that tail-based selection buys nothing over mean selection at the readable level, while research complexity consumes finite information budgets.

  • Problem

    Systematic trading needs a precise, testable account of when historical regularities persist under deployment rather than treating persistence as an axiom.

  • Method

    The framework declares latent-state and representation assumptions, prices their information costs, and uses split-sample tests to challenge the declared constants.

  • Results

    On the tested cross-sectional library, the tail criterion buys nothing over mean selection at the one level whose tail can be read.

  • Takeaways & Limitations

    Quantitative research should treat robustness and selection as budgeted purchases, because exceeding the data’s information budget can turn selection bias into apparent discovery.

  • Takeaways & Limitations

    The framework leaves representation choice unresolved, and noisy block scores can prevent even the intended penalty for Λ = 2 from being implemented.

Abstract

from arXiv · show

Systematic trading rests on one article of faith: that regularities found in the past persist. We state it as a time-invariant mechanism driven by an unobserved latent state, and show that it leaves a researcher five constants to declare --- the recurrence bound $Lambda$ at a block length $b$, the invariance defect $epsilon_0$ of the representation it is declared of, the coherence times $ell_i$ of the state's coordinates, the signal ceiling $rho$ and the fraction $kappa$ of it contingent on the regime --- after which the architecture of a correct quantitative investment system is nearly forced.

A Proofs … What is not an axiom.

The paper argues that a latent-state model with declared recurrence, representation, coherence, signal, and regime-contingency constants largely determines a correct quantitative investment architecture. It also separates axioms about the historically sampled world from deployment assumptions, showing that discovery can create an unvisited state with no finite historical guarantee.

  • 1 Introduction: A latent, time-invariant mechanism fails out of sample only when the future mixture over latent states moves beyond the past, so coverage—not induction itself—is the central problem.Bounding future state occupancy by Λ yields the system’s design constraints; without such a bound, the future may contain states absent from history.
  • 4 The axioms of latent regularity: The framework declares five governing quantities: recurrence bound Λ at scale b, invariance defect ε0, latent coherence times ℓ_i, signal ceiling ρ, and regime-contingent fraction κ.Λ controls occupation shifts, ε0 prices representation drift, ℓ_i charge persistence, ρ limits attainable signal, and κ sets ρ_blk = κρ; χ_σ is estimable rather than a declared constant.
  • A Proofs: The theory’s main results derive worst-case selection, information, robustness, and canonical-system rules from those constants rather than treating practitioner conventions as independent choices.Worst-case future risk is CVaR at level 1/Λ; capacity and search fit an information budget ρ^2 n_eff; and the canonical form proves five stages necessary under the axioms.
  • 2 Related work: The paper assembles established tools into a falsifiable architecture, while acknowledging that bounded losses, independence-model exhaustiveness, and axiom minimality remain explicit limitations.Its contribution is the assembly of density-ratio, mixing, CVaR, and information-budget machinery for latent-state markets, not originality of each component.
  • 4 The axioms of latent regularity: The representation must be declared at a level: coarsening lowers Λ but raises ε0, and the resulting future-risk bound pays an additive 2Mε0 penalty.This prevents vacuous existential representations and makes the mechanism–state boundary part of the model rather than a post hoc choice.
  • 4 The axioms of latent regularity: The framework distinguishes robust integrated coherence times from stronger exponential mixing assumptions, because rare long-lived regimes can evade mixing-rate summaries while still dominating a rule’s regime risk.The rare-ladder construction preserves a fixed exponential mixing rate but gives the relevant indicator a much larger coherence time, so ℓ_i must reflect the durations of regimes the rule depends on.
  • 4 The axioms of latent regularity: The signal ceiling ρ is simultaneously a per-period Sharpe ceiling and a bound on the best conditional-mean forecast’s population R^2, under the stated conditional-variance condition.The paper calibrates ρ^2 to roughly 10^-4–10^-2 per day, while κρ specifies the regime-contingent component relevant to deployability.
  • What is not an axiom.: The axioms do not guarantee that a newly discovered strategy remains historically covered: deploying capital creates a new latent coordinate, making the future occupation singular to the past for every finite Λ.Without an additional reflexivity assumption, no estimator based on history has even a weak guarantee for the deployed rule; reflexivity instead transfers risk through a specified two-parameter shift.

5 Why the axioms are necessary: an impossibility theorem

Theorem 5.1 shows that without sufficient sensitivity to replacing a 1/Λ fraction of the sample, any estimator can be forced toward the worst strategy while all axioms remain satisfied. Corollary 5.2 establishes that Λ is not recoverable from the observable past, making its choice an unavoidable domain judgment.

  • Theorem 5.1: An estimator insensitive to replacing ⌈n/Λ⌉ observations can be driven toward its class’s worst rule under a law satisfying the axioms.The adversarial construction preserves the axioms, changes at most a 1/Λ fraction of the past, and shifts the conditional mean by no more than Δ = 2ρσ̄.
  • Theorem 5.1: Finite-Λ guarantees therefore require large replace-⌈n/Λ⌉ sensitivity ε_n(⌈n/Λ⌉), because Axiom A3 forces the adversarial state to occupy that many visible periods.The replacement cannot be smaller than ⌈n/Λ⌉ periods, and at Λ ≥ n it can be a single period.
  • Selection statistics: Sample-mean ranking is vulnerable when its modal-choice margin exceeds 2M/Λ, whereas block CVaR_1/Λ with b ≤ k lies on the protected side of the theorem.The adversarial visit covers all but at most one block used by the block-CVaR statistic.
  • Corollary 5.2: Λ_n,m is not estimable from the past: conditionally it remains non-degenerate as n grows, and it equals ∞ with positive probability at every finite n.Consequently, no past-measurable estimator Λ̂_n converges to Λ_n,m in probability.
  • Corollary 5.2: Choosing Λ is therefore the researcher’s unavoidable domain judgment about unsampled future states, such as a currency peg breaking or clearing-house margin rules changing.The statistics can be automated, but the recurrence claim cannot be determined from the observable past.

6 The selection objective is derived, not chosen

Theorem 6.2 derives the model-selection criterion from the recurrence bound Λ: candidates must be ranked by a Λ-determined tail-risk functional rather than by a criterion chosen by taste. Its block and representation-scale consequences show how the protection changes with block length, dependence, and the information retained by the evaluation lattice.

  • Theorem 6.2: Given Λ, the optimal selection criterion is determined: it ranks candidates by the 1/Λ-tail mean of their block-risk law, with average backtest performance optimal only when Λ = 1.The bound is tight, equals the empirical mean at Λ = 1, and approaches the worst candidate risk as Λ grows without bound.
  • Mean–dispersion form: The resulting mean–dispersion coefficient depends on Λ alone in the Gaussian or long-block limit, while finite blocks retain dependence on the block-risk distribution’s shape.The Gaussian approximation has Berry–Esseen error O((ℓ/𝑏)1/2) relative to sdblk; at 𝑏 = 60, Experiment 7 finds the real block-score law is not yet Gaussian.
  • Block-scale reading: Longer blocks weaken the state-scale protection delivered by a fixed nominal Λ, so the criterion must account for coherence time ℓ and the block length 𝑏.At ℓ = 60 and nominal Λ = 4, the state-scale reading falls to 2.99 at 𝑏 = 60, 1.44 at 𝑏 = 693, and 1.26 at 𝑏 = 1452.
  • Representation lattice: The objective is a weighted sort: fill cells in decreasing order of estimated risk up to their Λ-weighted ceilings, yielding the 1/Λ-tail mean of the weighted cell law.Cell noise has conditional variance at most 𝐻σ2/𝑛𝑘, and dispersed cells can share less noise than contiguous blocks of equal size.
  • Answerability of the declaration: The declaration also supplies probabilistic protection: with probability at least 1 − ε, future deployment risk is bounded by CVaR1/Λ, while the pair of recurrence and invariance declarations is answerable even though their ratio is not.The corresponding generalization inequality holds with probability at least 1 − δ − ε, and the future-window law matches a historical window of the same length under Axiom A4.

7 The information budget

Section 7 defines an information budget from overlapping labels, persistent regimes, attainable signal strength, and adaptive search. It shows how this budget limits estimation and model selection, while breadth and regime-invariant signals determine where additional information can be gained.

  • Effective information: Overlapping labels and persistent regimes reduce effective information additively: overlap costs its full factor, while regime dependence costs ρ^2(2ℓ−1), not max(ℓ,H).At ρ^2=10^-2 and ℓ=60, the regime cost is 1.2, comparable to noise rather than sixty times larger.
  • Estimation ceiling: The attainable risk reduction is capped by ρ^2σ^2, and least-squares estimation reaches the corresponding upper rate up to relative O(ρ^2), with a matching lower bound.Persistence of observed features enters the contiguity scale and persistence term, but does not otherwise change the excess-risk rate.
  • Description length: The same budget has a description-length form: d effective parameters plus 2 log N nats, or d plus 2I(T;φ) for adaptive search, matching the PAC-Bayes divergence charge,.The selected candidate’s stationary risk exceeds its backtest by at most σ√(2(log N+log(2/δ))/n_eff) under Theorem 7.4(i).
  • Breadth and regime exposure: Breadth reduces noise and regime dispersion but cannot remove coherence or common regime variance; cross-sectional models therefore benefit more than timing models when signals are regime-invariant.Experiment 13 finds effective breadth an order of magnitude higher for published long-short signals, with a ceiling ratio of 16 versus breadth 21.
  • Representation choice: Representation choice cannot be recovered from the invariance axiom, so the proposed middle course is to declare several representations and pay the information cost of choosing among them.This extends the budget machinery to Stage S1, replacing both taste-based declaration and unrestricted selection.

8 The price of robustness … 1. An adapted, recurrence-improving feature map, with its constants declared. 𝑥𝑡= 𝜑(ℱ︀obs

The paper shows that robustness, search, training, and deployment draw on one finite information budget, then derives a necessary five-stage architecture whose constants constrain both attainable robustness and position sizing. It concludes that adapted representations, capacity control, contiguous-block evaluation, budgeted search, and conservative sizing are jointly necessary, while implementation choices within stages remain open.

  • 8.2 What the tail costs: Tail noise costs an effective-sample factor c(Λ)^2, with c(2)=1.17, c(4)=1.43, c(8)=1.78, and c(20)=2.47, but its price is nonmonotone across smooth and discrete regimes.For Gaussian block-risk laws, c(Λ) is approximately Λ^1/4; in rare discrete regimes it can equal Λ or collapse toward zero, so the applicable branch must be measured.
  • 8 The price of robustness: Robustness and search compete for one budget: selecting N candidates at recurrence level Λ requires 2c(Λ)^2 log N ≤ ρ^2 n_eff, making robustness unreachable when ρ^2 n_eff < 2 log N.When the constraint is feasible, it determines how much Λ the budget can buy; robustness is therefore a price, not a free safety margin.
  • 8.3 Robust training is a block re-weighting: Training on CVaR costs Λ_fit effective-sample units, whereas tail selection costs c(Λ)^2; when c(Λ)^2<Λ_fit, training on the mean and selecting on the tail is correct.In the scale-free isotropic case Λ_fit=Λ, and for Gaussian block-risk laws c(Λ)^2/Λ is 0.68, 0.51, 0.40, and 0.30 at Λ=2, 4, 8, and 20.
  • 8.4 The third cost: reading regimes through noisy blocks: Noisy blocks bias tail ranking upward: plug-in CVaR is not consistent at fixed block length, heteroskedastic noise favors candidates noisy in bad blocks, and longer blocks stop helping near b≈ℓ.At measured calibration, λ̄=0.674 below λ(2)=0.798, so no declared level implements even Λ=2’s intended penalty; the best attainable effective robustness is 1.6–1.7.
  • 9 Sizing: robust Kelly and the half-Kelly fixed point: The sizing corollary caps the Kelly fraction at one-half when budget utilisation κ=1, while full Kelly on an estimate has non-positive expected log growth.The optimal utilisation is generally interior: Experiment 2 finds it at every tested budget, with the resulting fraction between one-half and three-quarters.
  • 10 The canonical form: The canonical theorem makes all five stages necessary: recurrence-improving representation, capacity-bounded prediction, contiguous-block evaluation, budgeted search, and non-full-Kelly sizing each prevent a distinct quantified failure.Together the stages attain Theorem 8.1 for risk and Proposition 9.1 for sizing, but the axioms do not uniquely determine the implementation within a stage.
  • 1. An adapted, recurrence-improving feature map, with its constants declared. 𝑥𝑡= 𝜑(ℱ︀obs: The first architectural stage is an adapted feature map with declared recurrence scale, block length b, and invariance defect ε0, followed by a shrunk predictor whose effective dimension satisfies d_eff≲ρ^2n/H.Deployment as an ensemble raises the attainable fraction without consuming additional search budget, while the axioms constrain capacity rather than selecting ridge, boosting, or factor-model implementations.

11 What the theory is worth against practice · 12 Falsifiability and an empirical programme

The theory translates established quantitative-investment practices into an information-budget worksheet and specifies experiments that can test, refute, or qualify its declared constants. The empirical programme emphasizes measurable search costs, persistence-adjusted falsification, and the limits of validating Λ from finite histories.

  • 11.1 What the accumulated practice looks like from here: The paper presents independently developed field practices as consequences of its axioms, treating this correspondence as evidence that the axiomatisation captures the right structure.Table 4 is explicitly identified as the main evidence for this correspondence.
  • 11.2 Computing the budget: a worksheet: The information budget is computable largely from configuration using the label horizon H, effective dimension d_eff, and lifetime search size N.The worksheet defines d_eff through model-specific complexity measures and argues that researchers should record N, the number of evaluated proposals.
  • 11.2 Computing the budget: a worksheet: The worksheet treats ρ as a representation-level signal ceiling, uses declared ρ_blk to predict block-score autocorrelation, and requires reporting both declared and empirically refuted values.Theorem 7.4 determines the direction of in-sample error, while Experiment 7 tests whether ρ_blk equals ρ.
  • 11.2 Computing the budget: a worksheet: Experiment 7 estimates median c(Λ)^2 at approximately 0.9Λ across 71 daily timing rules, while Experiment 8 attributes the Gaussian excess to heteroskedastic score noise and associated bias.The estimate uses 30 years of one equity index with b = 60 and B = 124.
  • 11.3 What the experiments check: The thirteen-experiment programme keeps theoretical constants testable, with six experiments implementing the axioms literally so Λ, ε, the ℓ_i, ρ, κ, and d are known.All experiments reproduce from a fixed seed and are reported in Appendix B with their designs and findings.
  • 12 Falsifiability and an empirical programme: The programme predicts that CVaR selection beats mean selection by an advantage increasing with cross-block dispersion, while degradation scales with √d_eff H/n and √(2 log N υ/n).These predictions are falsified if the relevant slopes are not of order one or if persistence ℓ adds explanatory power beyond the stated terms.
  • 12 Falsifiability and an empirical programme: The tests can distinguish neither drift from an omitted stationary coordinate without time-scaling evidence nor a declaration from its search cost, so repeated held-out consultation must be treated as training data.Proposition 12.1 estimates search overdraft by comparing full-data search with a search truncated at the out-of-sample boundary, but the truncated arm is doubly handicapped.
  • 12 Falsifiability and an empirical programme: Falsifiability uses split-sample lower bounds and corrected dependence-aware tests: Λ cannot be estimated from above, but finite sampling imposes a minimum defensible floor from occupation variability.For p = 0.1 and ℓ = 60, the five-year floors are Λ_0.05 = 2.5 and Λ_0.01 = 3.1, rising to 3.4 for ten simultaneous cells at 95% confidence.

13 Limitations and open problems · 14 Conclusion

The paper concludes that its quantitative architecture is nearly forced by latent-state invariance, but remains limited by worst-case theory, declared quantities, representation choices, and omitted market effects. It proposes a priced, falsifiable workflow while identifying unresolved tests of the axioms, tail costs, and real-world deployment.

  • 13 Limitations and open problems: The framework remains limited by eighteen unresolved issues spanning worst-case bounds, a noisy-tail cost, axiomatisation, representation choice, and omitted market effects.Three issues are presented as decisive: which state representation is declared, the cost of reading regimes through noisy blocks, and whether the framework’s central tests can contradict that declaration.
  • 13.1 Worst-case bounds, and results proved for one chain: Heavy-tailed losses, one-chain derivations, insufficient block lengths, and path-dependent recurrence assumptions leave finite-sample guarantees incomplete.Theorem 8.1 assumes bounded block scores; several formulas are exact only for the redraw chain, while the mixing residual can be vacuous at the prescribed block length and recurrence can fail in a stationary world.
  • 13.2 The tail read through noisy blocks: The block-noise ranking bias does not vanish with more data, peaks near b≈2ℓ, and is only partly addressable through a costly narrowed level.The narrowed-level comparison exceeds budget by thirteen nats in Experiment 11 and performs worse on wide libraries because its tail averages one block rather than the required k.
  • 13.2 The tail read through noisy blocks: The tail-risk constants are not generally measured: c(Λ) is calibrated on block scores and self-designed rules, while Λfit depends on unestimated model quantities and need not be conservative.For raw block means, c2≈0.9Λ on twenty daily series and 82 long-only portfolios, whereas 133 published long-short signals sit at the Gaussian value; studentisation removes the excess.
  • 13.3 The status of the axiomatisation: The axiomatisation does not uniquely determine a representation or implementation: the ε(ψ)–Λ(ψ) trade is unsolved, the canonical form is a necessity theorem rather than a factorisation, and axiom minimality remains unproved.The signal ceiling is also declared at ρ but calibrated only at the weaker ρ1≤ρ, while heteroskedasticity is carried by a constant whose theorem-level treatment is incomplete.
  • 13.4 The market side: The market-side treatment is incomplete: reflexivity is reduced to one library-dependent number, transaction costs lack a budget-aware version, and at Λ=4 the tested real index produces no position.Experiment 13 measures publication decay at 0.76 of the in-sample mean and 0.44 afterward; Experiment 11 finds surviving edges only to Λ≈1.1–1.3, while transaction costs are c′=0.0035, or 3.5% of declared ρ.
  • 14 Conclusion: The conclusion reframes systematic trading as invariance conditional on an unobserved state, with Λ, ℓ, representation quality, signal strength, and regime dependence defining the practitioner’s information budget.The framework therefore prices robustness and recommends breadth, shorter horizons, better labels, pre-committed search, shrinkage, averaging, affordable tail ranking, and betting 1/(1+κ) of the unconstrained estimate.
  • 14 Conclusion: The framework is falsifiable: P2 did not fail across twenty series, but P1’s tail criterion added nothing over mean selection on the named cross-sectional library.The authors identify failure of P2—out-of-sample decay exceeding what label overlap and regime dispersion explain—as a reason to rebuild the budget, while acknowledging the contrary P1 result.

B Experiments

The experiments validate the paper’s dependence-adjusted selection, tail-risk, capacity, and information-budget claims in synthetic settings and across published and market datasets. They also identify when debiasing, invariance testing, and block-level training succeed or fail.

  • Experiments 1–6: −16.8 against +57.5 × 10−3 shows that mean selection can choose negative-future-return rules while block-CVaR remains positive at Λ ≥4.The optimal dimension follows ρ2n/H, with the exact optimum at κ = 0.33–0.94 of the ceiling; the PAC-Bayes bound is 1.11 times the Gaussian value.
  • Experiments 7–8: 1.95 versus 1.43 is the median ĉ(4) across 71 timing rules, while the invariance test rejects the term-spread declaration and preserves buy-and-hold performance only after correction.Debiasing works only at noise/signal ≤1; the observed series has ratio 6.6, and studentising removes the heteroskedastic score-noise excess over Gaussian.
  • Experiment 9: 19 of 20 assets have raw ĉ(4) above Gaussian, none do after studentising, and the excess decays as b−0.53; the term spread remains contradicted on only two corrected equities.The tail criterion’s excess degradation matches plug-in ĉ(Λ′) at both levels and is near the joint Gaussian null’s power boundary.
  • Experiment 11: A 72-way comparison is over budget, its haircut exceeds the selected rule’s edge, and S5 returns no position; an audit bounds the search overdraft at 0.86 annual Sharpe.Panel F confirms f⋆ = 1/(1 + κ).
  • Experiments 10 and 12: The fitted coefficient variance tracks H rather than ℓ⋆ = 60, while scaling slopes are 0.66 ± 0.06 for √dH/n and 0.79 for search, with ℓ adding nothing.The reported slopes are not refuted; the median Bartlett ratio is 0.86, with bond funds at 2–3 because of stale NAVs.
  • Experiment 13: 21.3 against 1.6 yields a ceiling ratio of 16 for 133 published long-short signals and 82 long-only portfolios, with long-only reproducing ĉ2 ≈0.9Λ and long-short matching Gaussian.P1 is a clean null at Λ = 2 with power, while post-publication decay is 0.44.

B.1 Experiment 1: the selection objective

In a two-state noisy backtest, mean-based selection ignores future-regime beliefs and turns negative under sufficiently large Λ, whereas contiguous block-CVaR remains positive. Shuffled-fold CVaR behaves like the mean and understates stress-regime variation by a factor matching the theoretical prediction.

  • B.1 Experiment 1: the selection objective: At Λ ≥ 4, mean-based selection produces negative expected future return, while block-CVaR remains firmly positive on the identical candidate set.Table 8 defines v as the selected strategy’s exact expected future return under a Λ-tilted future and identifies the oracle benchmark.
  • B.1 Experiment 1: the selection objective: Mean-based selection always chooses θ̄ = 0.561 regardless of Λ, revealing that average backtest performance has no parameter for beliefs about the future.The three contiguous criteria coincide exactly at Λ = 1, as Corollary 6.3 requires; their divergence becomes consequential at Λ = 4.
  • B.1 Experiment 1: the selection objective: Shuffled-fold CVaR selects θ̄ between 0.51 and 0.56 across Λ, versus 0.22–0.25 for contiguous blocks, and behaves like the mean.The stress-occupation variance is 0.0655 for contiguous blocks versus 0.0022 for shuffled folds, a 29.3 ratio versus the predicted 29.4.
  • B.1 Experiment 1: the selection objective: CVaR’s top-minus-median separation ranges from 23.4 to 55.7 × 10^-3 across Λ = 1.5–8, versus 20.0 × 10^-3 for the mean.The resulting ratio is 1.2 to 2.8, supporting the claim that the noise-only pricing in (8.2) is conservative when tail selection differs from mean selection.

B.2 Experiment 6: the price of the tail

Experiment 6 finds a third tail cost beyond robust training and selection: block-score noise creates a block-length-dependent ranking bias that can dominate daily calibrations. Narrower nominal levels help, but the attainable penalty is capped by the measured noise regime.

  • Panels A–B: For smooth block risks, c(Λ) tracks Λ^1/4 within 3% through Λ = 5 and 8% by Λ = 10, whereas rare discrete regimes can incur a nonmonotone Λ^2 cost.At Λ = 4, robust selection costs a factor of 2.0 in effective sample for smooth risks, so c(Λ) must be estimated from the application’s own block risk.
  • Panels C–D: At Λ = 4, CVaR selection returns 57.5 versus an oracle’s 80.0 at B = 315, while exact block signals recover 79.4, showing the shortfall is score noise.The deficit persists at B = 1260, where the selected value reaches only 60.3; the block-score noise is 0.158 against a signal of about 0.08.
  • Panel D3: The block-length sweep shows an interior optimum near twice coherence time: with ℓ = 40, value peaks at b = 80 at 60.2 but falls to 46.7 by b = 1280.The noise-to-signal ratio declines from 15.6 to 1.76, yet longer blocks destroy the dispersion the tail uses.
  • Panels E–G: Narrowing the nominal level from 1/Λ to 1/(3Λ) improves selection under both Gaussian and Student-t score noise, while studentising is dominated by raw CVaR.The recommended level is narrower by about three, not fifty, because stress blocks contain fragile candidates’ signal.
  • Panels C–J: The experiment identifies block-score noise as a third tail cost, alongside robust training’s Λ and robust selection’s c(Λ)^2, dominating the other costs at daily calibrations.The cost is charged in block length rather than sample size; narrowing the nominal level is the only effective lever.
  • Panel J: At measured noise, the sensitivity ceiling is 0.674 or 0.627 versus λ(2) = 0.798; nominal 1/4 achieves Λeff = 1.4, while the best level reaches only 1.6–1.7.At noise-to-signal 1, the λ(4) window opens at Λ′ = 12 in both noise families.

B.3 Experiment 5: robust training versus robust selection

Experiment 5 compares row-level with block-level CVaR training under two latent regimes and finds that robustness requires block aggregation: row-level CVaR nearly matches ERM, whereas block-level CVaR can turn losing rules positive above a budget threshold. Panel C also shows that Proposition 8.9 predicts training shortfall on a one-regime control without fitted parameters.

  • Experimental setup: The experiment uses two regimes with mixing time ℓ = 40, Gaussian features, and half invariant versus half sign-flipping coefficients; expected future Sharpe ratios are exact because ||w|| = 1.Linear per-block loss makes CVaR training convex, solved by projected subgradient descent on the unit sphere, with no evaluation noise.
  • Robust training versus robust selection: At Λ = 4, row-level CVaR reaches only 0.0001 versus 0.0348 for block-level CVaR in the ρ = 0.40 row, recovering almost none of its robustness.Row-level CVaR behaves like ERM on a smaller sample, supporting the claim that blocks are essential.
  • Robust selection: On a one-regime control, Proposition 8.9 predicts training shortfall closely without fitting: 0.840 versus 0.796 at Λ = 1 and 1.681 versus 1.595 at Λ = 2.Panel C uses d = 6, Λ ∈ {1, 2, 4, 8}, B ∈ [315, 5040], and 300,000 independent reference blocks.

B.4 Experiments 2–3: the arithmetic

Experiments 2–3 numerically validate the predicted effective-sample-size, feature-dimension, concentration, and coherence-time arithmetic across redraw, rare-ladder, and multiscale regimes. The results show that persistence and regime interactions—not nominal budget saturation alone—determine performance and variance.

  • Capacity: The exact optimal feature dimensions are d⋆=3, 8, and 45 for H=60, 10, and 1, while observed grid choices are 3, 5, and 20.At H=10, ℓ_x=250 retains d⋆=8, whereas ℓ_x=10 yields d⋆=5 because m falls to 0.65; half the loadings flipping in the rare state also gives d⋆=8.
  • Capacity: Budget utilisation remains below saturation at 0.94–0.33, while d=199 against n/ℓ_x=210 collapses out-of-sample Sharpe to 0.004–0.015.The collapse is attributed to the Marchenko–Pastur edge rather than the budget; interpreting the ceiling as the optimum would imply a Kelly fraction of 1/2, versus observed caps of 0.52–0.75.
  • Search: Measured nVar closely matches Lemma 7.1 across block correlations, coherence times, and label horizons, whereas charging every row by max(ℓ,H) would overstate variance.At ρ_blk=0, 0.1, and 0.3 with ℓ=60, measured values are 1.00, 2.21, and 11.2 versus 1.00, 2.18, and 11.6 predicted.
  • The rare ladder: The rare-ladder construction satisfies Axiom A4(ii) at declared ℓ⋆=5, with β0=2 to spare despite indicator coherence times of 19.6 and 40.0.Its exact linear-algebra diagnostics use ℓ⋆=5 and T=20, 40, 80, 160; the reported β(k)e^{k/ℓ⋆} peaks are 0.18, 0.09, 0.05, and 0.001.
  • Off the redraw chain: In the two-scale regime, additive rows reproduce the predicted variance sum to 4%, while an interaction at r12=0.3 gives nVar=2.59 versus 2.64 for its effective kernel and 43.7 for the slower coordinate.The results show that coincidence of two regimes ends when either does; charging the interaction at the slower coherence time was an upper bound by a factor of sixteen.

B.5 Experiment 7: the constants on one real daily series

Experiment 7 evaluates the widest constant, c(Λ), on a 1997–2026 daily equity-index series using 71 untuned timing rules. Real block scores place c2/Λ near 1 rather than Gaussian benchmarks, consistent with left-skewed, heavy-tailed scores.

  • Panel A: 0.9–1.1: real block scores give c2/Λ across levels, versus Gaussian benchmarks 0.68, 0.51, 0.40, and 0.30.Buy-and-hold gives ĉ(4) = 2.41 with a block-bootstrap 90% interval [1.80, 2.83], excluding the Gaussian 1.43.
  • Panel A: Median skewness −0.53 and excess kurtosis 1.92 accompany the shift in c from Λ1/4 toward Λ1/2.The passage attributes this shift to the score law’s left skew and heavy tails.
  • Panels H–H4: Across the real series, block-score autocorrelation is −0.01, coherence times are near 1.00, and prediction P1 is below the mean benchmark.The reported P1 values are 0.45 and 0.52 versus the mean’s 0.56, with t = −1.4 and −0.9.
  • Panel G: The alternative coordinate-cell procedure reduces sampling noise, but its declaration comparison changes with expressiveness and is not uniformly favorable.At matched cell counts, buy-and-hold retains +8.8 and +2.5 for the ten-cell term-spread declaration, versus +4.1 and 0.0 for nine joint cells.
  • Panel G2: Uncorrected testing rejects the term-spread declaration alone but not the volatility–term-spread pair; level correction withdraws that verdict.The paired declaration is additionally priced by expressiveness through the cell-mean comparison.
  • Panel J: Declared coordinates show power-law dependence over the first year, with exponents 0.35, 0.19, and 0.15, then plateau at 0.079, 0.110, and 0.209.The power-law fits have R2 values 0.99, 0.99, and 0.92; post-year values remain three to eight times the permutation floor.

B.6 Experiment 8: taking the measurement apart

Experiment 8 finds that measurement-error correction can recover latent-law functionals when noise is moderate, but information is unrecoverable at the observed high noise ratio. On real series, excess tail behavior is attributed to volatility clustering rather than latent-law non-Gaussianity.

  • Measurement-error limits: At noise-to-signal ratio 6.6, neither de-biasing estimator identifies the latent law: the naive plug-in reads 1.5 for every law, and B=1000 does not help.The mixture estimate is 0.98 ± 0.77 for a Gaussian whose c(4) is 1.43, indicating information loss in the convolution.
  • Measurement-error correction: At noise variance no greater than block-signal variance, the mixture recovers 3.25–3.32 versus 3.30 and 2.55–2.61 versus 2.66, while SIMEX closes about half the gap.De-biasing costs 2–6 blocks under these moderate-noise conditions.
  • Real-series diagnostics: On real series, noise correlates −0.26 with block scores and +0.50 with score magnitude, while studentisation reduces median excess kurtosis from 1.92 to 0.08.The plug-in remains at or below the Gaussian closed form at every level, and the reported excess over Gaussian values is attributed to volatility clustering.

B.7 Experiment 9: twenty assets, seven block lengths · B.8 Experiment 10: the effective sample, measured

Experiments 9–10 test the framework across twenty assets, seven block lengths, and the effective-sample adjustment, finding daily-data non-Gaussianity without persistent block-signal variance and limited dependence-based sample loss. The same experiments also show weak out-of-sample selection gains and largely rejected invariance declarations, while tail-selection degradation matches the predicted information-budget charge.

  • B.7 Experiment 9: twenty assets, seven block lengths: At b=60, candidate-median ĉ(4) exceeds Gaussian 1.43 by more than 0.1 on nineteen of twenty assets, but studentised values stay within 0.1 of or below 1.45 and excess kurtosis falls to −0.57–0.30.Raw medians range from 1.43–2.18, while the pre-studentisation range is 1.1–10.8 for pooled excess kurtosis.
  • B.7 Experiment 9: twenty assets, seven block lengths: Across block lengths from a week to a year, raw Gaussian excess decays as b^-0.53 and disappears at one year, while studentised values remain at or below reference and block-signal variance is unresolved.The variance ratio stays below one and flat, matching the Newey–West(10) long-run-variance floor; s^2 is below what thirty years can resolve.
  • B.7 Experiment 9: twenty assets, seven block lengths: For P1, the single best candidate reaches out-of-sample Sharpe 0.51 versus 0.18–0.39 for five robust criteria, whereas equal-weighted top-ten selection reaches 0.40 versus 0.45–0.53, with nothing clearing |t|=2.The comparison uses 752 equity candidates selected on expanding windows and re-selected yearly over 19 out-of-sample years.
  • B.7 Experiment 9: twenty assets, seven block lengths: The term-spread-only invariance declaration is contradicted on fourteen of sixteen equity series and all four bond funds, while evidence for the asset’s own stream remains limited to two equity positives and none for realised volatility.The equity term-spread defects are 0.012–0.042; the two matrix-priced bond funds show 0.008–0.068, attributed to NAV staleness rather than mechanism.
  • B.7 Experiment 9: twenty assets, seven block lengths: In the unfiltered 47-rule control, the tail criterion’s excess degradation relative to the mean is 1.97 at both tested levels, matching the plug-in prediction for Corollary 8.6’s budget charge.The ratio is computed after removing the library’s rule-average gap and is predicted using ĉ(Λ′) from the same training blocks.
  • B.8 Experiment 10: the effective sample, measured: Across Experiment 9’s twenty series and forty-seven rules, the median Bartlett long-run-variance ratio is 0.86 from one month to three years, far below the ℓ*=60 charge implied by n/ℓ*.Equity medians across rules are 0.72–1.00, while the high-yield fund reaches 3.3 because its matrix-priced NAV is autocorrelated by construction.

B.9 Experiment 11: the five stages walked once

Experiment 11 walks S1–S5 in order on committed data with constants declared first and records the information budget stage by stage. The resulting audit finds that the search and haircut budgets eliminate any supported position despite affordable fitting.

  • Setup: The protocol uses Experiment 7’s 71 rules plus one fitted candidate across precommitted training, embargo, and held-out eras, with the constants declared before computation.The experiment tests only whether S1–S5 can be executed sequentially and whether §7’s budget becomes a ledger with one line per stage.
  • S1: declared constants: Declared constants give υ=2.19, an effective sample of 2299, a 50.3-nat fit budget, and a 23.0-nat search budget, requiring at least 523 trading rules.The declared setup uses Λ=4, ℓ=60, ρ2=0.01, H=1, and τμ=1.
  • S5: decision: At the selected rule’s turnover of 0.117 per day, the haircut alone explains the negative outcome, while buy-and-hold has no position under any reading.The selected rule has c′=0.0035; its edge survives only to Λ=1.3 or 1.2 under two readings, whereas buy-and-hold reaches only about Λ≈1.1.
  • S5: decision: Every candidate’s robust edge disappears at the self-consistent floor, with all edges crossing zero between Λ=1.1 and 1.3 below Λ̂split=1.36.Figure 4 marks the declared Λ=4 and shows that the held-out-window floors at 2.24 and 1.17 eliminate every candidate.
  • What the walk shows: The 72-way search is over budget, while the haircut alone permits fewer than 4.4 proposals, so no block length produces a position.At measured dependence, the search budget is 45.5 nats and the Gaussian charge fits with 9.5 nats to spare, but the plug-in charge does not.

B.11 Experiment 13: the cross-sectional library

Experiment 13 tests the missing cross-sectional dimension using two libraries with similar macro exposure but sharply different effective breadth. Greater breadth reduces regime noise and improves ceiling bounds, yet selection deficits persist and corrections broadly agree.

  • Panel A, breadth: Effective breadth was 21.3 versus 1.6, while average pairwise correlation was 0.052 versus 0.660 across the long-short and long-only libraries.The breadth estimates were stable to within 10% across three windows.
  • Panel A, breadth: The long-short library’s ceiling ratio was 16.4 versus 29 from breadth alone, while averaging reduced the regime term tenfold, from 0.37 to 0.037.The model-free block-mean variance ratio gave 11.0, compared with the proposition’s 16.4 ceiling ratio.
  • Panel B, the plug-in: Studentisation made both libraries Gaussian at Λ = 4, with values 1.39 and 1.40, although the unstudentised long-only median ĉ(4) was 1.88 versus 1.43 Gaussian.For the long-only library, c2/Λ was 0.94, 0.88, 0.80, 0.53 at Λ = 2, 4, 8, 20; the long-short values were 0.63, 0.53, 0.42, 0.26.
  • Panel C, P1: P1 selection underperformed mean selection on the long-short library by −0.46 at 1/4 and −0.73 at 1/8, with t = −2.4 and −3.3.The protocol used expanding annual re-selection and evaluated both selected-candidate Sharpe and pooled lower-quarter tails.
  • Panels D–F: In the blocked split, degradation was 0.269 out of 0.716 for long-short versus 0.052 out of 0.194 for long-only, and the deficit persisted in the real-time library.The real-time deficits were −0.21 at 1/4, −0.33 at 1/8 (t = −2.0), and −0.35 at 1/12 (t = −2.2); deflated Sharpe, PBO, and SPA, agreed with the correction verdicts.

B.12 Experiment 14: the same five stages, where the answer is known · B.13 The five predictions: power, and what happened · C Notation, and the symbols that are used twice

Experiment 14 shows that the five-stage procedure rejects null and zero-edge decoys while recovering known signal when blocks are long enough, with conservative declaration choices causing measurable refusals. The prediction tests find limited power for several claims, a substantially smaller effective than nominal complexity, and one prediction left unconfirmed; the notation consolidates the paper’s symbols and conventions.

  • B.12 Experiment 14: the same five stages, where the answer is known: 93.0% of null-world draws have a positive best backtest mean, yet the procedure returns no position in all 1200 walks after the S4 haircut.The max-of-72 statistic has median 0.0177 and maximum 0.0565 against a 0.0412 haircut; the remaining 3.2% is closed at b = 60.
  • B.12 Experiment 14: the same five stages, where the answer is known: Sensitivity costs block length rather than sample size: nothing clears S5 at b = 60, while four fifths of draws do at b = 1008 with the rule declared alone.With 72 decoys, S3 still selects the invariant rule, but conservative ρblk = ρ and block-tail construction cause additional refusals and coverage failures.
  • B.12 Experiment 14: the same five stages, where the answer is known: The robust edge is below the deployed tilted-regime mean in 93.5% of 4369 opened positions, with median margin +0.053 versus the +0.041 promise, and no zero-edge decoy opens.Invariant-rule coverage is 99.5% over 3494 positions; deployment tilts stress-regime weight from 0.1 to 0.4.
  • B.13 The five predictions: power, and what happened: P3’s measured effective breadth is 21.3 versus nominal 133 for long-short signals and 1.6 versus 82 for long-only portfolios, while P4 tracks levels but not cross-library ordering.The asserted P3 ceiling ratio is 16 versus 1.7 across the libraries, approximately tenfold; P4’s ratios are 1.97 on ĉ(4) = 1.82 and 2.92 on ĉ(12) = 2.82.
  • B.13 The five predictions: power, and what happened: P5 remains unconfirmed because κ requires the forecast’s shrinkage factor and is informative only when identified more tightly than Experiment 2’s 0.52–0.75 interior-optimum band.The prediction is therefore not treated as already confirmed.
  • C Notation, and the symbols that are used twice: The notation defines Λ as the recurrence bound, ℓ as coherence time, ρ as the signal ceiling, ρblk as regime dispersion, ε0 as representation defect, and κ through ρblk = κρ.It also records b as evaluation scale, ℓ* = max(ℓ, H, Lx) as the contiguity scale, and T’s mutual information as at most log N for a fixed list.
  • C Notation, and the symbols that are used twice: The notation distinguishes representation, block-score, capacity, and tail quantities, warns that six letters are reused by context, and notes that loss and reward tail conventions yield the same c(Λ).Key definitions include Ψ’s menu and selection charge, ΔN and ΔCVaR selection errors, κ̌ = d/(ρ^2 nfit), and the Kelly fraction s^2/(s^2 + τ^2).
Loading 2608.23416v1…