Source-linked AI summary
Free-Probability Kernels for Zero-Rollout Hyperparameter Selection in Reservoir Computing
Sara Malacarne, Andrea Ceni, Claudio Gallicchio
TL;DR
Reservoir hyperparameter selection usually requires repeated finite-reservoir rollouts, despite reservoir computing’s lightweight readout. The paper uses free-probability cross-lag moments to construct a deterministic large-width kernel for ranking candidates from labelled pilot data. Across benchmarks, this selection matches exhaustive references closely while substantially reducing rollout cost.
Problem
Finite-reservoir hyperparameter selection repeatedly instantiates reservoirs and evaluates validation error, with cost depending on candidate grids, reservoir widths, and random realizations.
Method
A deterministic large-width kernel approximates finite-reservoir feature geometry and ranks candidate operating points from a short labelled pilot sequence without finite-reservoir rollouts.
Results
Across ten temporal benchmarks, FP shows no significant difference from task-direct simulation-based selectors requiring 156 600 selection rollouts, while FP-K matches the full-grid reference using 4.8% of its rollout budget.
Takeaways & Limitations
The results support free-probability kernels as deterministic surrogates for selecting reservoir operating regimes when validation rollouts are scarce.
Takeaways & Limitations
The exact theory requires linear recurrent dynamics, and pilot–deployment distribution shift has not been studied.
Abstract
from arXiv · showhide
Reservoir computing (RC) couples a fixed recurrent dynamical system with a trained lightweight readout, but this efficiency is partly lost during hyperparameter selection: the recurrent gain, input scale, and leakage rate determine the reservoir's stability and temporal processing regime and are usually tuned through many rollouts. We introduce a deterministic, pilot-informed selector for leaky linear reservoirs followed by coordinate-wise nonlinear features. Free probability yields cross-lag propagation coefficients that summarize how the reservoir mixes past inputs. In the large-width limit, these coefficients define a deterministic temporal kernel that approximates the finite-reservoir feature geometry. Kernel ridge regression on a short labelled pilot sequence therefore ranks candidate operating regimes without instantiating or rolling out a reservoir, and the selected configuration transfers across widths. Across ten synthetic temporal benchmarks, zero-rollout selection obtains a mean deployment score of $0.772$, compared with $0.774$ for exhaustive simulation-based search, while avoiding $156\,600$ selection rollouts. With a small rollout budget, the proposed ranking provides the strongest mean performance at every tested budget and reaches the exhaustive reference using $4.8\%$ of its rollout cost. On four public electricity-transformer-temperature (ETT) forecasting datasets, five retained candidates recover the exhaustive operating point on three datasets. On multivariate cellular-traffic forecasting, 15 rollouts per cell reach the 462-rollout exhaustive reference and outperform random search and Bayesian optimization at low budgets. These results position free-probability kernels as deterministic surrogates for selecting reservoir operating regimes when validation rollouts are scarce.
I. INTRODUCTION
Reservoir hyperparameter selection is costly because recurrent gain, input scale, and leakage shape task-dependent dynamics and require repeated finite-reservoir validation. The paper replaces these rollouts with a pilot-informed deterministic kernel that approximates finite-reservoir feature geometry for candidate ranking.
- Motivation: Reservoir computing trains only a lightweight readout over a fixed recurrent representation, but performance depends strongly on recurrent gain, input scale, and leakage rate.These hyperparameters govern stability, memory, and the relative influence of recent and distant inputs.
- Motivation: Standard selection repeatedly instantiates reservoirs, generates trajectories, trains readouts, and evaluates validation error across candidate grids, widths, and random realizations.The selected configuration may depend on search width, requiring repeated selection when deployment width changes.
- Approach: The proposed selector evaluates candidates with a deterministic large-width kernel on a short labelled pilot sequence, without instantiating or rolling out candidate reservoirs.The same procedure can also pre-screen the candidate grid before a small number of finite-reservoir evaluations.
- Approach: Free probability computes mixed cross-lag propagation moments for leaky linear reservoirs, whose nonlinear feature Gram matrix converges at large width to a deterministic temporal kernel.The kernel is constructed for each operating point θ = (σr, σin, α) and approximates the corresponding finite-reservoir feature geometry.
- Novelty: The paper uses large-width kernels for task-informed hyperparameter selection rather than solely for reservoir characterization or prediction.Kernel ridge regression ranks candidates on labelled pilot data before the selected configuration is deployed in a finite reservoir.
- Novelty: The analysis addresses temporal dependence caused by repeatedly applying the same recurrent matrix and establishes deterministic nonlinear feature similarities at large width.This is presented as the first task-informed zero-rollout method for selecting finite-reservoir hyperparameters with such a kernel.
A. Leaky linear reservoir and stability
The model uses a leaky linear recurrence with fixed random recurrent and input matrices, followed by coordinate-wise nonlinear features and a trained readout. Its stability and pilot-kernel notation are defined through the recurrence, candidate parameters, feature blocks, and kernel-ridge prediction.
- Model and stability: The reservoir uses a scaled real-Ginibre recurrent matrix and an independent Gaussian input matrix, with the recurrent gain and input scale controlling their normalization.The asymptotic analysis takes reservoir width n → ∞ while keeping input dimension fixed.
- Model and stability: For linear recurrence, the echo-state property is equivalent to the propagated influence of initial conditions vanishing as A^t → 0.This characterizes when trajectories driven by the same inputs forget different initial states.
- Model and stability: For α > 0, the asymptotic linear echo-state condition ρFP < 1 is equivalent to σr < 1.The finite-width analysis additionally introduces a robustness margin.
- Readout: The fixed coordinate-wise feature map is zt = ψ(xt), with experiments using ψ = tanh and theory assuming a bounded continuously differentiable activation with bounded derivative.The recurrent and input matrices remain fixed while only the output readout is fitted.
- Readout: Pilot training and validation indices define kernel blocks indexed by time steps, and kernel ridge regression predicts validation outputs from the training block.The finite-reservoir kernels are formed from untruncated trajectories, while the theory first derives finite-context limits and then the complete-history limit in the stable regime.
C. Finite-context free-probabilistic kernel construction
The finite-context construction replaces random reservoir feature geometry with deterministic quantities derived from cross-lag propagation and limiting state covariances. Coordinate-level self-averaging and concentration yield a nonlinear kernel that supports pilot-based selection.
- Finite-context construction: Starting from x0 = 0, the linear recurrence is unrolled and truncated to the most recent L + 1 propagated inputs for finite-context selection.The resulting construction retains a fixed context length before extending to complete washed-out history.
- Propagation coefficients: Mixed normalized traces of propagated recurrent matrices determine the state covariance because the recurrence is linear.These cross-lag coefficients quantify correlations between input components injected at different lags.
- Propagation coefficients: The Ginibre cross-lag propagation coefficients are derived by expanding mixed powers, applying circular-element limits, and controlling fluctuations by Gaussian concentration.The derivation supplies the deterministic propagation profile used by the kernel.
- Self-averaging: Diagonal propagation terms self-average across coordinates, making the random diagonal profile deterministic despite individual reservoir units remaining random.This coordinate-level result is required because the nonlinear activation acts neuronwise.
- Nonlinear kernel: The deterministic linear-state covariance is passed through the coordinate-wise activation to obtain the nonlinear feature kernel.For tanh, the required Gaussian expectation can be evaluated by deterministic low-dimensional Gauss–Hermite quadrature.
- Nonlinear kernel: For fixed context length and bounded deterministic inputs, the empirical nonlinear kernel converges in probability to the deterministic large-width kernel.On a fixed pilot set, the resulting empirical Gram blocks converge to the deterministic temporal feature geometry.
D. Complete history and other recurrent ensembles
The construction extends the finite-context kernel to complete-history limits and other recurrent ensembles, while implementing candidate ranking through pilot-indexed deterministic kernel computations. Under stability and a unique minimizer, this ranking is consistent with finite-reservoir selection as width grows, subject to stated limitations.
- Complete history: When σr < 1, increasing context length L recovers the washed-out reservoir state with geometrically controlled error.The stable recurrence forgets remote inputs geometrically, linking finite-context selection to the complete-history limit.
- Other recurrent ensembles: The same kernel construction extends to Haar-orthogonal, cyclic, circulant, skew-symmetric, and complex-valued recurrent ensembles with suitable propagation moments.The experiments themselves use real Ginibre recurrence.
- Zero-rollout selector: The kernel functions as a virtual reservoir, capturing the large-width temporal feature geometry of each finite reservoir without constructing recurrent weights or state trajectories.Candidate-specific kernels are computed directly from the short labelled pilot sequence.
- Zero-rollout selector: Algorithm 1 ranks stability-guarded candidates using deterministic kernel ridge validation on pilot train and validation indices.The procedure takes a pilot sequence, context length L, candidate grid Θ, and ridge grid Λ, then returns θ⋆ and a ranking.
- Zero-rollout selector: No n × n recurrent matrix is formed, and the selection computation uses pilot-time indexing rather than reservoir coordinates.Its computational costs do not depend on deployment width n.
- Consistency: With fixed context and finite candidate and ridge grids, a unique deterministic validation minimizer is selected by finite reservoirs with probability tending to one as n →∞.The result follows from kernel convergence, continuity of kernel ridge regression, and finite-grid minimizer stability.
- Limitations: The consistency result concerns agreement on the pilot validation objective, not minimization of deployment test error.Nearly tied candidates may require larger context or reservoir width; the synthetic erf-versus-tanh setting also falls outside the corollary verbatim.
IV. EXPERIMENTS
The experiments test zero-rollout and budgeted selectors across synthetic, public ETT, and cellular-traffic forecasting tasks under matched candidate grids and deployment protocols. They compare deterministic FP ranking with exhaustive, random, and Bayesian-optimization alternatives while controlling regularization, sampling, and rollout budgets.
- Synthetic benchmarks: Ten synthetic benchmarks cover nonlinear temporal processing, explicit memory, delayed nonlinear transformations, and chaotic-system forecasting.The suite includes NARMA-k, memory capacity, delayed input, Lorenz63, Mackey–Glass, and Lorenz96 tasks.
- Real-data benchmarks: The real-data evaluation uses four ETT-small datasets and a proprietary cellular-traffic benchmark spanning hourly and 15-minute forecasting settings.ETT predicts future oil temperature, while Telco jointly predicts two traffic KPIs across ten cells.
- Selectors: FP ranks candidates by deterministic kernel NMSE on labelled pilot data without instantiating a finite reservoir.Synthetic FP selection uses three independent pilots and assigns each candidate its largest validation NMSE across them.
- Budgeted selectors: FP-K applies Direct500 only to the top K FP-ranked candidates, whereas Random-K and TPE-K use the same retained candidate budget for comparison.Each task-direct trial uses three selection seeds, so a budget K costs 3K finite-reservoir rollouts; FP pre-ranking adds none.
- Kernel implementation: Synthetic FP selection uses an erf surrogate, while real-data FP selection uses the exact tanh kernel evaluated by Gauss–Hermite quadrature.The erf surrogate is restricted to the controlled synthetic setting and is separately checked against the exact kernel.
- Synthetic protocol: The synthetic Cartesian grid contains 6 480 raw points, with 5 220 admissible candidates after the fixed stability guard.Exhaustive selection therefore uses 15 660 rollouts per task and 156 600 across ten tasks, while FP uses zero finite-reservoir rollouts.
- Forecasting protocol: Forecasting selects one configuration per dataset or Telco cell using a validation score aggregated over all forecast horizons.A single readout jointly predicts all horizons from one reservoir state, and every selector optimizes the same aggregate objective.
- Real-data protocol: ETT and Telco grids each contain 245 raw candidates, of which 154 satisfy the stability guard.Exhaustive empirical selection requires 462 finite-reservoir rollouts per ETT dataset or Telco cell.
B. Synthetic benchmark results
Across ten synthetic tasks, FP delivers deployment performance close to exhaustive simulation-based selection while eliminating finite-reservoir search rollouts. Its ranking also remains competitive under statistical comparison, width changes, and limited rollout budgets, though performance is not uniformly monotone across finite widths.
- Zero-rollout selection: 0.772 mean deployment score at n = 20 000 places FP close to Direct500’s 0.774 while using zero finite-reservoir search rollouts.FP is best or tied-best on MC, NARMA30, NARMA50, and Lorenz96, with larger gaps on Lorenz63 and Inubushi.
- Average-rank comparison: 2.05 average rank places FP behind Direct500 at 1.65 but ahead of ESN Direct500 at 2.90 and Memory500 at 3.40.After Holm correction at α = 0.05, FP and Direct500 outperform Memory500 significantly, while differences among FP, Direct500, and ESN Direct500 are not significant.
- Width portability and selection cost: FP selection cost is independent of reservoir width, whereas empirical selectors’ measured cost increases with nselect because they instantiate and evaluate finite reservoirs.FP timing varies across repeated measurements, but nselect is not an input to the selector.
- Width portability and selection cost: 0.692 at n = 1000 rises to 0.772 at n = 20 000 for the same FP operating point, narrowing its gap to Direct500 from 0.028 to 0.002.The same operating point is reused without rerunning selection; the reported interpretation is consistent with convergence toward the deterministic kernel, not guaranteed monotonicity at every finite width.
- Width portability and selection cost: FP is stable or improving across reported deployment widths on every task, whereas selectors tuned at nselect = 500 can show non-monotone trends on individual tasks.The width trend is presented as consistent with the large-width construction, but not as monotone improvement at every finite width.
- Budgeted screening: FP-K gives the strongest mean performance at every tested rollout budget, with its advantage largest at small K and shrinking as K grows.At K = 250, FP-K reaches 0.777 versus 0.774 for exhaustive Direct500 using 750 finite-reservoir selection rollouts per task, or 4.8% of the full-grid cost.
- Selection-length sensitivity: FP reaches 0.771 by Tselect = 500 and is close to its highest peak already at Tselect = 250, while direct selection continues benefiting from longer sequences.Memory500 remains constant because its task-agnostic proxy selects the same candidate at each tested sequence length.
C. Real-data benchmark results
Real-data experiments show that FP screening is most effective when rollout budgets are scarce: it recovers exhaustive operating points on several ETT datasets and nearly matches or exceeds exhaustive selection on Telco with far fewer rollouts. The method also supports width-independent selection, while its exact theory remains limited to linear recurrent dynamics and matched pilot–deployment distributions.
- Zero-rollout selection: On ETTm1, zero-rollout FP scores 0.505 versus 0.499 for full-grid Direct500, while ESN Direct500 reaches 0.508.The task-agnostic Memory500 proxy improves the longest-horizon ETTm2 scores but lowers the 1-hour score, and remains below task-informed selectors on dataset averages.
- Budgeted screening: FP-K recovers the full-grid Direct500 operating point at K = 5 on ETTh1, ETTh2, and ETTm2.On ETTh2, Random-K and TPE-K remain below FP-K even at K = 50; on ETTm2, they approach FP-K only at larger budgets.
- Budgeted screening: 0.622±0.112 matches or slightly exceeds the 0.621 ± 0.113 full-grid Direct500 reference on Telco using 15 rather than 462 rollouts per cell.At K = 5, FP-K also exceeds Random-K and TPE-K, which score 0.548 and 0.572 respectively.
- Budgeted screening: FP-K beats Random-K and TPE-K on all ten Telco cells at K = 5, K = 10, and K = 25, with adjusted pHolm ≤0.0078.FP-50 recovers the full-grid Direct500 operating point with 150 rather than 462 rollouts per cell; differences are no longer significant by K = 50.
- Width-independent tuning of reservoir dynamics: The large-width kernel makes selection cost independent of deployment width, allowing the same θ⋆ selected with Tpilot = 500 to transfer across widths without rerunning search.FP plateaus at Tselect = 500, whereas task-direct selection continues to improve.
- Scope and extensions: The exact theory requires linear recurrent dynamics, and behavior under pilot–deployment distribution shift has not been studied.The method also ranks a fixed discrete candidate grid rather than refining candidates continuously.
APPENDIX A FIXED-CONTEXT KERNEL
The fixed-context analysis derives deterministic cross-lag coefficients for repeated recurrent propagation and uses them to characterize the nonlinear feature kernel at large width.
- Fixed-context kernel: The proof shows that mixed propagation traces converge to deterministic coefficients τ_k,ℓ.These coefficients summarize cross-lag effects generated by repeated use of the same recurrent matrix.
- Fixed-context kernel: The averaged diagonal propagation terms converge to the same deterministic coefficients across reservoir coordinates.Gaussian Poincaré control gives Var(m_n(i,j)) = O(n^-2) for the mixed trace quantities.
- Nonlinear feature-kernel limit: Conditional on the recurrent matrix, each coordinate pair is Gaussian, so its limiting covariance identifies the nonlinear feature kernel.The argument then uses concentration over the Gaussian input matrix to remove finite-width fluctuations.
- Technical control: For bounded deterministic inputs, every fixed power of the recurrent matrix has uniformly bounded norm moments.The expected norm is at most 2 + o(1), with a sub-Gaussian upper tail.
C. Diagonal self-averaging
Diagonal self-averaging establishes that coordinate-level covariance terms converge to deterministic limits, enabling the nonlinear kernel limit after fluctuation control.
- C. Diagonal self-averaging: Permutation invariance makes the diagonal entries identically distributed, allowing their average to represent the diagonal propagation behavior.The recurrent matrix law is preserved by simultaneous row and column permutation.
- C. Diagonal self-averaging: Gaussian Poincaré bounds give Var((B_n)₁₁) = O(n^-1), yielding diagonal self-averaging after a finite union bound.The bound controls all retained lag pairs for fixed L.
- C. Diagonal self-averaging: Only diagonal propagation terms are needed for the marginal Gaussian law of each coordinate pair entering the kernel.Dependence between distinct reservoir coordinates is handled separately by concentration.
- Nonlinear feature-kernel limit: The nonlinear covariance map is defined by h(Σ) = E[ψ(g₁)ψ(g₂)] for a Gaussian pair with covariance Σ.Continuity of this map transfers covariance convergence to nonlinear feature-kernel convergence.
- Nonlinear feature-kernel limit: Conditional Gaussian concentration makes fluctuations around the conditional mean vanish in probability at rate O_P(n^-1).On any fixed finite pilot set, this gives entrywise and Frobenius-norm convergence of kernel blocks.
APPENDIX B COMPLETE HISTORY AND SELECTION CONSISTENCY
The complete-history analysis shows that, under stability and bounded inputs, finite-context kernels converge as the retained lag grows, with uniform control over the history tail.
- Selection consistency: The finite-context deterministic kernel converges uniformly over the finite candidate grid and pilot-kernel entries used by the selector.This follows from the covariance-tail control and continuity of the nonlinear Gaussian covariance map.
- Complete history: Under σ_r < 1 and bounded inputs, the complete-history covariance series converges absolutely.The nonnegative cross-lag coefficients and their generating function provide the convergence argument.
- Complete history: The omitted lags outside {0, …, L}² are bounded through the union of the regions k > L and ℓ > L.Symmetry reduces the tail estimate to a one-sided bound applied to covariance entries.
- Complete history: A geometric Ginibre power bound controls recurrent-history tails uniformly for sufficiently large reservoir widths.For 1 < R < σ_r^-1, q_R = 1 − α + ασ_rR < 1 supplies the geometric decay factor.
C. Complete-history kernel
The complete-history kernel converges under the stable assumptions, and deterministic selection is asymptotically consistent with finite-reservoir selection on finite grids when the optimum is unique.
- C. Complete-history kernel: Theorem 3 establishes convergence of the Ginibre kernel to the complete-history kernel for every fixed time pair.The proof combines the fixed-context result with geometric finite-width and context-tail bounds.
- Selection consistency: With matching feature maps, a unique deterministic validation-score minimizer is selected by finite reservoirs with probability tending to one as width grows.The result assumes finite candidate and ridge grids and strictly positive ridge parameters.
- Selection consistency: In the stable regime, a unique complete-history optimum persists for all sufficiently large finite contexts.Selection based on the complete finite-reservoir trajectory converges to this optimum as width tends to infinity.
- Selection consistency: Near ties can require larger reservoir width or maximum lag before the candidate ranking stabilizes.The issue arises because a small score gap leaves less tolerance for approximation error.
- Haar-orthogonal extension: The same deterministic-kernel conclusions extend to Haar-orthogonal recurrent matrices, including the complete-history result when σ_r < 1.Orthogonal invariance replaces the Gaussian concentration argument used for Ginibre matrices.
APPENDIX D OTHER RECURRENT MATRIX STRUCTURES
The appendix extends the kernel construction to several recurrent-matrix structures by substituting structure-specific cross-lag coefficients into the common covariance and nonlinear-kernel formulas. Ginibre, Haar-orthogonal, and cyclic-shift matrices share fixed-context coefficients, while their complete-history controls differ.
- Common coefficient framework: The kernel derivation depends on recurrent-matrix structure through limiting cross-lag coefficients and diagonal self-averaging properties.Once these coefficients are known, the covariance and nonlinear kernel follow from the common construction.
- Closed-form surrogate: A closed-form kernel surrogate is used only for synthetic selection and timing experiments, while ETT and Telco results use the exact tanh kernel with Gauss–Hermite quadrature.The surrogate's corresponding limiting kernel entries differ by less than 2δerf < 0.071.
- Recurrent-matrix catalogue: The real Ginibre, scaled Haar-orthogonal, and cyclic-shift structures produce the same fixed-context coefficients.The cited proposition identifies ηk,ℓ with τk,ℓ for the real Ginibre case and extends the same coefficients to Haar-orthogonal matrices; the appendix also states this for cyclic shifts.
- Recurrent-matrix catalogue: Complete-history control differs across the three structures: Ginibre uses a proposition, whereas orthogonal matrices admit a direct geometric bound.Thus matching fixed-context coefficients does not imply identical control arguments for full histories.
C. Gaussian skew-symmetric matrices
This section discusses Gaussian skew-symmetric and related recurrent structures through self-averaging and complete-history conditions, while noting that complex-valued extensions require additional activation conventions. The surrounding experiments report selection-length sensitivity and validate the synthetic surrogate.
- Gaussian skew-symmetric matrices: For Gaussian skew-symmetric matrices, permutation symmetry and a Gaussian-Poincaré argument provide diagonal self-averaging, while normality supports complete-history control under an asymptotic condition.The supplied passage states these properties but does not include the condition's displayed formula.
- Complex-valued matrices: Complex recurrent matrices replace transpose with the Hermitian adjoint, and nonlinear feature kernels require a specified complex activation or realification convention.The experiments use real-valued recurrent matrices.
- Deployment evaluation: Table VII reports 1 − NRMSE mean±STD across deployment seeds for five widths, with best and second-best scores marked per task and width.At n = 20 000, scores use ten deployment seeds; intermediate widths use three.
- Selection-length sensitivity: Selection-length sweeps report deployment performance at n = 20 000 and per-task selection CPU time as functions of Tselect at nselect = 500.These results are summarized in Table VIII and Figures 5–6.
- Surrogate validation: The erf-surrogate and exact-tanh selections differ by at most 6.3 × 10−3 per task and by O(10−3) on average across checked selection lengths.The paper reports these differences as below task-level variability in the deployment tables.