Source-linked AI summary

REFLEX: Reflexive Equilibrium Fixed-point Learning for Endogenous eXchanges

Vignesh Nagarajan, Shriraghav Ashok

arXiv:2608.16155v1cs.LGcs.CEcs.GT

TL;DR

Dealer retraining can destabilize OTC bond markets because quotes reshape the trading flow used for subsequent learning. REFLEX estimates this feedback loop from measurable market-making features and finds that competition amplifies instability, while structurally anchored correction stabilizes cases where blind retraining collapses.

  • Problem

    Existing stability theory uses constants that desks cannot measure before deployment, limiting quantitative assessment of whether quote retraining and dealer competition remain stable.

  • Method

    REFLEX derives the retraining constants from microstructure primitives in a structural OTC market-making model and verifies the resulting stability extensions in a learned simulator loop.

  • Results

    Competition amplifies instability by effective dealer count, while structurally anchored correction stabilizes loops where blind retraining collapses.

  • Takeaways & Limitations

    Algorithmic liquidity-provision stability should be assessed as a market-level property, with structural anchoring preferred over additional free-form model capacity.

  • Takeaways & Limitations

    Calibration is proxy-level rather than trade-level TRACE and omits per-dealer inventories, state-dependent performativity, and richer policy classes.

Abstract

from arXiv · show

In over-the-counter corporate bond markets, dealers compete for client trades by quoting bid and ask prices. Tighter quotes attract more business, but also informed customers more likely to trade ahead of adverse price moves, leaving the dealer holding the risk. As dealers increasingly use machine learning to set quotes, they retrain these models on the trades their own quotes attract, creating a feedback loop in which each model reshapes the market that generates its next training data. The question is therefore not only whether a quoting model performs well, but whether the market it creates stays stable as the model learns from it. Existing performative prediction theory gives a sharp stability condition, yet expresses it through abstract properties of the learning objective a trading desk cannot measure before deployment. We introduce REFLEX, a framework that replaces those unobservable quantities with three measurable features of dealer behavior: how strongly trading volume responds to tighter quotes, how sharply the dealer's objective bends around its optimum, and how quickly informed flow increases as spreads narrow. REFLEX combines these into a single retraining modulus, a pre-deployment stability margin estimated from a desk's own quote and execution history that predicts whether repeated retraining will converge or amplify itself. In simulation, predicted and measured stability agree within 8%, and competing dealers increase instability by 1.74x with two and 3.16x with three, as predicted. Where ordinary retraining becomes unstable at modulus 1.21, a structurally anchored correction converges as blind retraining collapses. Calibrated over 36 years of public market data, stability headroom falls roughly 4.4x for investment grade and 4.3x for high yield from calm to crisis regimes. Ultimately, REFLEX turns an abstract convergence theorem into a market-level safety margin.

1 Introduction

REFLEX addresses the instability of repeated retraining in OTC bond quoting, where dealer policies alter the flow distribution used for future training. It replaces deployment-inaccessible theorem constants with measurable structural quantities and validates stability, correction, competition, robustness, scaling, and cadence predictions.

  • Motivation: Tighter quotes attract more volume and more toxic informed flow, so each quoting policy induces the distribution on which it is next trained.Repeated fitting, redeployment, and refitting constitute repeated risk minimization in performative prediction.
  • Motivation: The condition ε < γ/β is sharp, but desks cannot evaluate its Lipschitz constants before deployment or assess competition-driven fragility from post-trade spreads alone.The same measurement problem applies to supervisors observing multiple dealers retraining against a shared informed-flow pool.
  • Core framework: REFLEX derives γ, β, and ε from fill-curve curvature, P&L scale, and toxic-flow slope, making m = εβ/γ available before deployment; measured moduli track predictions within 8% where the loop contracts.The framework embeds performative prediction in a GLFT-family structural OTC market-making model and checks closed forms against a learned simulator loop.
  • Stability correction: The closed-form PerfGD correction converges past the RRM boundary, while its structurally anchored learned version reaches the realized performative optimum where blind retraining collapses.The correction uses one extra scalar per step.
  • Multi-dealer stability: 1.74× and 3.16× amplification at N = 2, 3 demonstrate the multi-dealer prediction that competition destabilizes the market by Neff before any single dealer would.The multi-dealer boundary is ε < γ/(Neffβ).

2 Related Work

Prior work has developed performative prediction theory across increasingly rich settings and derived market-making models with fixed flow responses. REFLEX addresses the gap between these literatures by computing performative constants from market structure.

  • Performative prediction: Performative prediction work established convergence for RRM and stochastic variants, then extended to response estimation, dynamic distributions, bandit feedback, multi-agent games, and robust formulations.Across this literature, (ε, β, γ) remain abstract constants of an unspecified …
  • Market-making models: Market-making research derived optimal quotes under exponential fill intensities and later incorporated multi-asset reduction, adverse selection, stochastic control, and flow–price interaction.These models treat the flow response to quoting as a fixed input.
  • Research gap: No prior work computes performative constants from market structure.REFLEX is positioned to fill this gap by linking the constants to measurable market behavior.

3 Market Model and the Retraining Loop

REFLEX models each dealer deployment as quoting at a half-spread, with benign fills following an exponential curve and informed flow governed by a spread-gated channel. Because the learner freezes the deployed toxic level while retraining on candidate spreads, retraining follows a cobweb map that converges to the performatively stable spread if and only if m < 1.

  • Market model: Each deployment is one period of dealer quoting at half-spread h, with zero skew in the scalar theory and a multi-bond extension.The scalar theory uses latent liquidity and separates benign and informed notional channels.
  • Market model: Benign notional follows a GLFT exponential fill curve, while informed notional follows a spread-gated channel parameterized by arrival, elasticity, intensity, informativeness, and toxicity-feedback terms.Informed flow incurs adverse-selection severity ψ > 0 per unit.
  • Retraining loop: The learner is structurally blind because each deployment freezes toxic level T = τ(h_dep), allowing only the benign channel to vary with candidate spread h.The dealer therefore best-responds to the deployed regime rather than fully updating toxicity during retraining.
  • Retraining loop: Retraining forms a cobweb map that contracts to the performatively stable spread h_SP if and only if m < 1.This is the stated convergence condition for the frozen closed-form model.
  • Simulation: The simulator extends the closed forms with multi-bond flow, informed-flow saturation, and liquidity inflation driven by realized flow.These endogenous channels are omitted from the frozen closed forms and exposed by subsequent measurements.

4 Closed-Form Stability Theory

The section derives closed-form stability boundaries for retraining loops from measurable curvature, volume-response, and toxic-flow terms, with stability iff the retraining modulus m<1. It extends the theory to corrected objectives, competing dealers, robust estimation, multi-bond scaling, and lazy deployment.

  • R1 (analytic boundary): Stability holds if and only if the retraining modulus m<1, with every term evaluable before deployment.The boundary follows from the implicit-function theorem and predicts loop behavior before retraining begins.
  • Structural corollaries: Widening spreads reduce toxic-flow sensitivity exponentially, so the modulus saturates at the self-consistent operating point and may remain below 1 under default-like constants.Measured instability therefore describes the retraining map at the operating spread.
  • R2 (PerfGD correction): The PerfGD correction converges where m>1 because performative-objective curvature remains positive beyond the blind retraining boundary.The correction adds the distribution-response term Δ to the blind gradient and costs one scalar more per step.
  • R3 (multi-dealer systemic risk): Competition creates systemic instability when the common-mode eigenvalue reaches its critical dealer count, even though each dealer individually satisfies m1<1.Shared toxic spillover synchronizes dealers through a rank-one coupling, making fragility a market property.
  • R5 (factor scaling): For d bonds, stability is governed by the spectral-radius condition ρ(M)<1, while Woodbury reduces k-factor computation to O(dk^2).Truncation error is linear in the residual factor variance λk+1(C).
  • R6 (lazy deployment): For m>1, lazy deployment remains stable when K≤Kmax, and a deadbeat cadence Kdb produces one-shot convergence.Warm-started K-step deployment interpolates between inertia and the exact cobweb, while fitting the plain boundary can underestimate effective curvature.

5 Experimental Setup

The experimental setup closes a calibrated multi-bond OTC market simulator with four retraining modes and estimates stability through independent response-sensitivity instruments. Protocol audits, public-data calibration, provenance limits, and numerical proof certificates define how results are measured and interpreted.

  • Market simulator: The simulator models benign and informed clients, inventory, capped informed flow, latent liquidity, and a learned market-response operator fit over past deployments.With window ≥2, the operator identifies learned dD/dϕ from deployment-policy summaries.
  • Retraining modes: Four modes close the retraining loop: Blind RRM, PerfGD-analytic, PerfGD-learned, and structurally anchored PerfGD-structural.PerfGD-structural uses fitted response families, a 12-deployment history, 35% relative-step trust region, and anti-echo freezing.
  • Stability measurement: Three instruments estimate ε: common-random-number best-response probes, optimal-transport Wasserstein sensitivity, and fitted informed-flow curves.The CRN probe uses paired deployments at href ± δ; the transport estimate uses debiased log-domain Sinkhorn divergences, and the toxic-flow curve differentiates fitted spread-response relation (1).
  • Protocol audit: 0.05 collection jitter is retained because 0.2 inflates the BR probe ∼3×; results use medians and IQR over 8 seeds with R4 robust bands.The gain f, rather than adversariality α, is swept because α produces a non-monotone response of 0.08 → 1.83 → 0.67.
  • Data calibration and provenance: ∼36 years of daily and ∼70 years of monthly public series calibrate configurations by rating × volatility regime alongside TRACE-derived factors and returns for 212 real-CUSIP bonds.Only (A,k,σ,h) are data-identified; the toxic channel is structurally scaled, and crisis intensity is flagged as degenerate at k = 0, n = 74 days.
  • Verification: 66 proof certificates numerically re-derive the load-bearing identities, inequalities, and dynamics on raw and calibrated real-unit configurations.The setup also formalizes the logical skeletons of all six results in Lean 4.

6 Results

Results validate REFLEX across numerical certificates, regime-level market data, retraining dynamics, multi-dealer amplification, and robustness checks. The findings show that stability headroom contracts sharply in crises, predicted boundaries track measured behavior, and structural or cadence controls can prevent instability.

  • Certificates: All 66 numerical certificates pass, with worst residuals inside tolerance; the calibrated pass validates the identities in real units.The excluded beyond-boundary raw-unit cell is deliberate because applying its constants to per-$100-par calibrated units would violate the stated conventions.
  • Market-data results: 4.4× and 4.3×: stability headroom collapses from calm to crisis for IG and HY, respectively, while HY remains more than 10× below IG in every regime.The index reaches a crisis plateau because of the degenerate crisis fit, so intra-crisis ranking is not identified; headroom spans 465 → 2.2 from IG-calm to HY-crisis.
  • Falsification test: 8%: measured modulus is 0.390 versus predicted 0.426 at f = 2, while the measured crossing is f* ≈ 3.17 versus the a-priori 4.70.The deployment’s own flow inflates the liquidity field to ρ ≈ 2.3, making the analytic boundary a conservative lower anchor rather than an unbiased point prediction.
  • Multi-dealer market: 1.74× and 3.16×: competing dealers amplify instability at N = 2 and N = 3, matching predicted amplification of 2 and 3 within 13% and 5%.The critical dealer count is Nc ≈ 7.9, and synchronized common-mode dynamics—not differential spillover—drive the instability.
  • Stability correction: 1.21: in the f = 6 demo regime, blind retraining diverges while the structurally corrected one-dimensional ascent converges.Blind RRM ends at half-spread 0.16, whereas learned unconstrained corrections fail to identify the toxic slope off the deployed regime.
  • Retraining cadence: K ≲ 6: lazy retraining keeps the RRM-unstable market measurably stable, but divergence returns beyond the stability window.At f = 5, the measured anchor is m̂ = 1.205 > 1; the predicted exit Kmax = 5.75 aligns with measured |μ| = 0.695 at K = 5 and 1.5 at K = 8.

7 Conclusion

REFLEX makes performative-prediction stability desk-evaluable from market microstructure primitives and validates its extensions through learned-loop experiments, real-market fragility analysis, and machine-checked identities. The conclusion emphasizes structural anchoring over additional model capacity while documenting calibration, protocol, and verification limitations.

  • Contributions: REFLEX computes the stability tuple (𝜀, 𝛽,𝛾) from microstructure primitives and validates extensions through predict-then-verify learned-loop tests.The extensions cover correction, competition, finite samples, dimension, and retraining cadence.
  • Contributions: 36 years of real market data support REFLEX as a daily fragility index for evaluating market-level retraining stability.The paper positions stability as a market-level rather than solely desk-level problem.
  • Contributions: Structural anchoring stabilizes the learned corrected loop where blind retraining collapses, whereas a free-form alternative fails despite using the same data and more capacity.The conclusion recommends spending modeling budget on structural anchoring rather than network capacity.
  • Limitations: Calibration remains proxy-level rather than trade-level, using VIX-implied spreads, no per-dealer inventories, a structurally scaled toxic channel, and a degenerate crisis cell (𝑘= 0).Regime ordering is data-driven, but absolute critical gains are not.
  • Limitations and future work: 66 numerical certificates provide the verification of record because the Lean 4 skeletons were reviewed but not yet compiled.Future work includes trade-level calibration, inventory-state dynamics, richer policy classes, and compiling the formal layer.

Reproducibility

REFLEX is fully reproducible through an openly available framework containing its derivations, software, calibration pipeline, verification layer, Lean sources, and paper-grade reports. All paper results can be regenerated deterministically with one command in approximately 25 CPU-minutes.

  • Reproducibility: The release includes derivation documents D1–D6, the simulator, operator, four loop modes, estimator suite, calibration pipeline, data catalogue, 66-certificate verification layer, Lean sources, and illustrated reports.These components collectively cover the framework, calibration, verification, formalization, and reported runs.
  • Reproducibility: ∼25 CPU-minutes is sufficient to regenerate all paper results with one deterministic command, run_all –profile full, from the configuration and seed.The regeneration procedure is deterministic when run with the specified configuration and seed.
Loading 2608.16155v1…