Source-linked AI summary

VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery

Yuqi Li, Siyuan Liu, Bingjun Liu

arXiv:2609.07065v1cs.AI

TL;DR

Agent-to-agent alpha discovery is slowed by repeated feedback cycles and communication that is not stably replayable. VST structures those exchanges for multi-cycle prediction and transactional verification, finding that typing preserves accuracy while adding auditability and leakage control; its descriptive holdout results remain limited by benchmark underperformance, selection, costs, and validation scope.

  • Problem

    Repeated feedback cycles dominate long searches, while contemporary A2A messages lack a stable replayable contract for predicting and safely skipping interaction.

  • Method

    VST wraps communication in typed, causally addressable records and uses a multi-cycle predictor with transactional verification and rollback.

  • Results

    An equal-information free-text channel reaches the same predictor hit rate, showing structured transport adds correctness properties rather than predictive accuracy.

  • Takeaways & Limitations

    VST provides a governable, replayable pipeline for verifiable rollback-safe decisions, not a demonstrated return or raw-speed advantage.

  • Takeaways & Limitations

    Evaluation covers one CSI 1000 universe with gross returns and descriptive leap counts, while selected Sharpe and development-calibrated gates limit interpretation.

Abstract

from arXiv · show

Agent-to-agent (A2A) alpha discovery is slowed by repeated feedback cycles between mining and evaluation agents, whose hand-offs, in contemporary LLM multi-agent systems, are free-form natural-language messages that carry no stable contract and cannot be replayed. We first restructure this communication as a structured agent-to-agent protocol of \emph{typed, causally addressable, unicast records}, so that the committed stream forms a causal trajectory. On that trajectory a single predictor with four typed heads forecasts the accumulated guidance the two miners would receive several cycles ahead; a transactional verify--leap controller then commits a multi-cycle speculative outcome only when it passes a four-level gate, and otherwise rolls back to the exact prior state. Structure is the enabling contribution, and its value is not accuracy. A controlled ablation shows an equal-information free-text channel reaches the same predictor hit rate. What typing provides is a state that can be schema-checked, replayed deterministically, and prevented by construction from leaking a forecast to an evaluator: auditability by construction, not an empirically stress-tested guarantee. On a CSI~1000 out-of-sample holdout, our single run is the only one among eight methods (seven baselines and ours) to hold a positive median annualized return and Sharpe at the factor level, though the median return \emph{in excess} of the benchmark stays negative for every method including ours; its development-selected top-20 portfolios reach a $0.71$ median holdout Sharpe, selected on a split inside the optimization horizon. We report these single-run results descriptively, gross of costs, and are explicit about their limits throughout; in particular we do not isolate the effect of the leap machinery from the inherited search substrate, which we leave to future work.

1 Introduction

VST treats structured communication as the prerequisite for forecasting and verifiably skipping multi-agent feedback cycles. Its contribution is auditability and controlled state transitions, not improved prediction accuracy.

  • Repeated miner–evaluator–report round-trips dominate wall-clock cost in long evolutionary searches.
  • Contemporary A2A systems primarily use free-form or loosely structured messages that lack a stable, replayable contract and causal trajectory.
  • VST wraps transmissions as typed, causally addressable, replayable records with routing, visibility, and transactional provenance.
  • A single predictor forecasts multi-cycle miner guidance, while a transactional controller verifies and atomically commits or rolls back speculative outcomes.
  • An equal-information free-text channel reaches the same predictor hit rate, so typing supplies correctness properties rather than an accuracy gain.
  • The protocol preserves separated predictor and evaluator skill evolution so successful forecasts do not teach evaluators to agree with them.
  • A verify-then-commit unit spans a multi-cycle transition of two coupled loops rather than a single generated token.

2 Related Work

The paper places VST at the intersection of structured LLM-agent communication, trajectory prediction, alpha discovery, and speculative execution. Its distinction is a transactional, visibility-typed communication record for multi-cycle agent workflows.

  • LLM-agent systems commonly coordinate through natural language, while whole-interaction contracts remain less structured than individual hand-offs.
  • VST extends typed communication envelopes with transactional provenance and visibility typing for replayable, leak-resistant, rollback-safe speculation.
  • Unlike single-agent trajectory policies, VST predicts a multi-cycle guidance state and receives transaction-level feedback from coupled agent loops.
  • Alpha-discovery methods generally optimize factor quality as a monolithic search, whereas VST targets the communication layer around an inherited search substrate.
  • VST applies verify-then-commit to a multi-cycle transition of two coupled loops, rather than to tokens or memory words.

3.1 System Setting and Problem Formulation

VST wraps an inherited dual-loop alpha-search substrate and formulates acceleration as predicting and validating future committed states. Evaluation separates development, test, and fully isolated holdout roles while constraining leap quality against an unaccelerated baseline.

  • 3.1.1 Dual-loop system substrate: The dual-loop topology, six evaluators, and evaluator metric and skill libraries are inherited unchanged from prior work.
  • 3.1.1 Dual-loop system substrate: Four temporal splits distinguish training, validation, test, and holdout data, with only holdout untouched by every closed-loop decision.
  • 3.1.1 Dual-loop system substrate: The test segment remains inside the optimization horizon because meta-review agents can inspect its outcomes, whereas holdout is fully isolated.
  • 3.1.2 Objective and quality constraint: The objective is to predict a future miner-guidance state, advance both loops beyond one cycle, validate actual outputs, and skip at least two ordinary cycles when successful.
  • 3.1.2 Objective and quality constraint: The accelerated state must match the unaccelerated baseline optimum within tolerance ϵ under the same splits, budget, and stopping rule.
  • 3.1.2 Objective and quality constraint: The baseline reference is computed only on development partitions, and failed trials revert state evolution to baseline with speculative overhead.

3.2 Structured Inter-Agent Communication Protocol

The structured communication plane wraps existing round-JSON payloads in typed records that encode identity, routing, causality, contracts, and provenance. Deterministic unicast and visibility policies make the resulting trajectory replayable and prevent speculative forecasts from reaching evaluators.

  • The protocol adds a canonical routing wrapper without replacing the existing operational round-JSON format or receiver API.
  • Packets are A2A-compatible, causally addressable, deterministic, unicast, and visibility typed while preserving producer, consumer, and state-transition semantics.
  • Its universal envelope contains identity, routing, causality, and contract groups, while provenance links packets to source artifacts and transaction status.
  • The same fields that clarify ordinary transmission also provide stable coordinates for prediction over the communication trajectory.
  • The gateway materializes one physical packet per destination and assigns role-specific body contracts to miners, evaluators, and report agents.
  • Visibility validation admits forecast packets only on permitted miner routes and rejects speculative packets addressed to evaluators before persistence.
  • The structured record is the load-bearing contribution enabling trajectory forecasting, verifiable leap commitment, and predictor–evaluator separation.
  • Figure 1 combines the three-layer wrapper, native payload fields, deterministic routes, and retained prediction targets.

3.3 Multi-Horizon Structured Trajectory Prediction

Once committed hand-offs form a typed, replayable trajectory, one shared predictor forecasts multi-cycle miner guidance through four typed heads and emits schema-constrained speculative bundles. The targets aggregate future structured guidance directly across horizons rather than predicting future factors or chaining one-step forecasts.

  • Trajectory and targets: The committed trajectory contains visibility-legal packets among the report agent, miners, evaluators, and controller, with miner guidance extracted from four exact A2A field paths.The two report-agent channels and two evaluator-aggregate channels are aligned by state.
  • Trajectory and targets: For each horizon k ∈ K = {2, 4, . . . , Kmax}, the target is accumulated structured guidance delivered to both miners over the next k ordinary cycles.The target is guidance, not a future factor, metric, verdict, or score.
  • Trajectory and targets: The type-aware aggregator composes set edits, resolves repeated promote/demote decisions by recency and support, and retains the latest bounded categorical adjustment.Aggregation dispatches by field type rather than applying one reducer to every channel.
  • Predictor: The predictor reads a sliding trajectory window and requested horizon k, then emits a schema-constrained bundle, calibrated success confidence, and epistemic uncertainty.Its loss combines type-appropriate field terms, schema validity, and a Brier term for realized commit outcome.
  • Predictor: Four channels share one global representation and four output heads, while the communication layer delivers two speculative unicast packets directly to the respective miners.The design uses one predictor rather than four independent agents or policies.
  • Training corpus: Replay reconstructs recorded tokens and states rather than re-executing nondeterministic agents, and rollback traces remain outside the committed supervised trajectory.Training pairs one causal prefix with direct multi-horizon targets rather than chained one-step predictions.

3.4 Transactional Verification and Leap Control

The controller treats a forecast as a speculative transaction: real miners and evaluators execute it in scratch state, and a four-level validation gate either commits the complete transition or resumes normally after rejection. The design uses conservative consistency checks and explicitly bounded transactional semantics.

  • Validation pipeline: A predicted bundle is executed by both miners and six evaluators in a scratch namespace, and validation concerns the speculative outcome rather than field-by-field forecast matching.The four levels short-circuit in ascending order of cost.
  • Adaptive horizon selection: The controller selects the farthest admissible horizon with k ≥ 2, sufficient confidence, bounded uncertainty, and predicted saved compute; otherwise it performs no leap.A trial counts as a speedup only when coupled refinements b are fewer than the replaced cycles k.
  • L1 bundle validation: Before speculative execution, the gate rejects malformed fields, stale parents, short horizons, and evaluator-targeted predictions.This provides schema and visibility checks before miners and evaluators run.
  • L2 cross-loop consistency: Cross-loop consistency requires executable factors, legal evaluation configuration, admissible metrics, and portfolio references supported by the alpha-side output.The enumerated referential checks are necessary conditions found sufficient for observed divergences, not a complete characterization.
  • L2 cross-loop consistency: The consistency screen is conservative rather than a soundness guarantee because conflicts outside the enumerated checks could be admitted.Any detected violation fails the level because downstream quality scores would be uninterpretable on an incoherent joint state.
  • Validation pipeline: Figure 3 checks the bundle, actual miner outputs, blinded evaluator reports, and development evidence; only an all-pass path commits atomically.Failure discards the scratch transaction and resumes one normal cycle from the current committed state.
  • L3-L4 validation: Evaluator risk aggregation preserves disagreement by taking the worst severity across three independent reports instead of averaging away a critical finding.The empirical development gate recomputes factor and portfolio quality on two development partitions while isolating final test and holdout data.
  • Commit and rollback semantics: The transactional model is single-writer and strictly serial, with one speculative transaction rooted at the current committed state and one controller coordinator.Atomic commit excludes crash atomicity and durability, while exact rollback is best-effort over tracked application stores.

3.5 Separated Co-Evolution of Predictor and Evaluator Skills

VST separates predictor and evaluator skill evolution so the predictor learns from transactional outcomes while evaluators judge only actual outputs. This preserves forecast independence and prevents speculative failures from silently changing active evaluation standards.

  • Predictor skill evolution: Commit or rollback outcomes provide outcome-grounded feedback for predictor improvement, with skills reinforced only when actual execution passes L1–L4 and saves cycles.Recurring high-loss feedback patterns can also spawn specialized trajectory skills.
  • Predictor skill evolution: The predictor stores reusable skills for feedback aggregation, cross-loop consistency, trajectory detection, horizon calibration, and rollback diagnosis.Its policy selects and composes these skills while adapting to changing mining regimes.
  • Predictor skill evolution: The predictor uses one shared policy with four typed heads to model a coupled transactional state rather than coordinating separate RL agents.Supervised learning fits bundle content, while reinforcement learning fits horizon and utility.
  • Evaluator skill evolution: Evaluators receive only actual candidates, outputs, metrics, portfolio results, and assigned historical views; predicted identifiers, confidence, uncertainty, and fields are excluded.This prevents forecast anchoring while allowing evaluator rubric and meta-rubric development from observed outputs.
  • Separation and safety: Separating skill libraries prevents the predictor from teaching evaluators to agree with its forecasts, preserving L3 independence and blocking reward hacking.Transactional promotion also prevents speculative failures from silently changing active evaluation standards.

4 Experiments

Experiments compare VST with seven baselines on matched CSI 1000 holdout evaluations and test whether its structured transport and search process generalize beyond development splits. VST is the only method with positive median absolute holdout return and Sharpe, while selected portfolios show stronger gross results but face selection and cost limitations.

  • Out-of-sample comparison: +0.059 median annualized return was the only positive holdout result among the eight methods.Every baseline had negative median annualized return, including AlphaGen at −0.001, AlphaMCTS at −0.015, and Alpha-GPT at −0.031.
  • Out-of-sample comparison: −0.028 median excess return versus the CSI 1000 benchmark remained negative for VST and every other method.The result supports a positive absolute-return comparison, not benchmark outperformance.
  • Portfolio-level performance: 0.71 median holdout Sharpe was achieved by the development-selected top-20 portfolios, alongside +13.3% median annualized return.The selected set had 80% positive portfolios and 45% above Sharpe one; selection used test-split Sharpe and holdout was not used for selection.
  • Selection concerns: 1.36 was the estimated chance maximum Sharpe on the development scale across 852 candidate portfolios.The selected twenty portfolios had development Sharpe between 1.50 and 1.98, but this benchmark is explicitly a selection-only bound.
  • Robustness across splits: 2.234 to 0.346 was VST’s median Sharpe decline from validation to holdout, while every baseline crossed into negative holdout territory.The authors describe this as a robustness effect, not a training-fit effect, and do not attribute it to leap machinery alone.

5 Discussion

VST’s structured transport makes speculative multi-cycle execution auditable and rollback-safe, while the ablation finds no predictor-accuracy advantage over equal-information free text. Its broader applicability depends on predictable guidance and verification that costs less than the cycles replaced.

  • Typed, causally addressed records make speculative leaps schema-checkable, replayable, and structurally unable to leak forecasts to evaluators.Verification commits only valid speculative outcomes; otherwise the system rolls back to the exact prior state.
  • An equal-information free-text channel reaches the same predictor hit rate, so structured transport is not claimed to improve prediction.The paper frames typing as a correctness and auditability mechanism rather than an accuracy source.
  • Multi-agent searches can benefit when guidance is low-dimensional and slowly varying enough to predict several cycles ahead, and verification is cheaper than executing those cycles.The paper identifies verification cost as the binding condition for realizing savings.
  • Failed speculation incurs speculative and verification overhead without committed progress when no horizon is admissible or every trial rolls back.The quality constraint is treated as an empirically estimated bound with thresholds fixed before final testing.
  • The selected top-20 portfolios have a +0.71 median holdout Sharpe but a higher +0.87 mean, indicating a positive, heavy-tailed distribution.The portfolio set was selected during development, so the median is presented as the more conservative summary.

6 Limitations

The evaluation is limited to one market and gross returns, while gate calibration, speculative efficiency, and safety claims remain bounded by validation and construction-based evidence.

  • The evaluation uses only CSI 1000, so transfer across markets or asset classes is untested.Transferability is argued in principle rather than demonstrated empirically.
  • Reported return and risk-adjusted figures are gross and exclude transaction costs, slippage, and turnover constraints.The cross-method comparison is internally consistent, but absolute magnitudes are not net of costs.
  • Development-estimated gate thresholds and horizon choices may admit low-quality states within tolerance; rollback ensures state safety, not quality optimality.Highly non-stationary trajectories can also reduce admissible horizons and prevent recovery of speculative overhead.
  • Leakage prevention, exact rollback, and predictor/evaluator separation are established structurally rather than through injected-leakage or documented failure-mode stress tests.The paper therefore characterizes auditability as a construction property, not an empirically stress-tested guarantee.

7 Conclusion

VST turns a dual-loop alpha-discovery communication layer into a substrate for verifiable multi-cycle acceleration, with structured transport enabling prediction, gated commitment, rollback, and separated agent skills.

  • VST’s C1 is a three-layer, A2A-compatible transport with deterministic unicast routing and visibility typing that makes transmissions typed and replayable.It preserves existing payload paths and receiver arguments while serving as the architectural prerequisite for downstream components.
  • A single four-headed predictor forecasts accumulated miner-directed guidance over a chosen horizon on the committed trajectory.The predictor operates after the communication layer has been structured.
  • A four-level transactional gate commits valid speculative outcomes in one logical step or rolls them back to an exact prior state.Separated skill libraries keep predictor and evaluator evolution independent of forecast agreement.

A Aggregation Algorithm

The type-aware aggregator combines structured guidance according to each field’s schema, preserving cycle order, bounded numeric values, and the latest committed text.

  • Set-valued fields accumulate additions and removals in cycle order, collapsing repeated decisions using recency and supporting-cycle counts.This applies to fields such as accepted metric sets and active factor tags.
  • Numeric and bounded-categorical fields retain the latest value clipped to the schema’s declared range.
  • String advice fields retain the most recent committed text.

B Operational Runtime Side Effects

Operational effects were measured from closed-loop per-round records under a shared machine and period. Structured transport showed lower median wall-clock time, while success-rate differences were small and mixed across portfolio- and factor-level measures.

  • Success-rate effects were small and mixed, with portfolio-level measures favoring structured transport but the single-factor evaluator pass rate favoring free text.
  • 353 s versus 385 s: the structured arm had a lower median per-round wall-clock time than the free-text arm.The reported interquartile ranges were [335, 373] and [363, 414], respectively.
  • The wall-clock comparison is observational because both arms shared a machine and period rather than a controlled timing experiment.
  • 0.991 versus 0.985: the structured arm had a slightly higher fraction of rounds whose composite passed the structure gate.
  • 2.28 versus 2.24: the structured arm had a slightly higher median composite Sharpe at the portfolio level.
  • 0.246 versus 0.128: the free-text arm had the higher per-candidate evaluator pass rate at the single-factor level.
Loading 2609.07065v1…