Source-linked AI summary
EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale
Ege C. Kaya, Abolfazl Hashemi
TL;DR
EGGROLL’s low-rank perturbations improve scalability but can change the mean optimization field despite identity covariance. The paper characterizes this effect and introduces LOO-ROLL, which improves equal-time post-training outcomes without significant losses.
Problem
Dense perturbations make moving large populations through language models computationally difficult, while low rank may alter the mean optimization dynamics.
Method
The paper characterizes EGGROLL’s population field as a resolvent-filtered gradient and introduces LOO-ROLL, a leave-one-out estimator using one evaluation per direction.
Results
At matched wall time, LOO-ROLL improves seven of ten tested settings in individual paired tests, with no significant loss; five gains survive Holm correction.
Takeaways & Limitations
Rank-one EGGROLL can retain nearly all of dense ES’s sampling efficiency despite differing mean dynamics, while LOO-ROLL uses the evaluation budget more effectively.
Takeaways & Limitations
The variance comparison relies on a local affine model, and the experiments cover finite objectives and models up to 8B parameters.
Abstract
from arXiv · showhide
EGGROLL makes evolution strategies (ES) practical for LLMs by replacing dense Gaussian weight perturbations with low-rank Gaussian products, often of rank one. This choice is computationally attractive but geometrically severe: each rank-one perturbation lies in a zero-volume subset of the ambient matrix space, despite having identity covariance. We characterize the mean EGGROLL update field at finite rank and nonzero perturbation radii, then analyze the error of its finite-population estimator. The population field is obtained by applying an explicit resolvent to the gradient of the objective smoothed by the perturbations. We show that the resolvent can introduce a nonconservative component and can reverse the local stability of an optimum. EGGROLL is nevertheless exact on every quadratic objective at every rank and radius. For smooth objectives, its first local finite-rank correction is $O(σ^2/r)$, and nonasymptotic bounds control the resulting field error under smoothness assumptions. Under a local affine model, rank-one perturbations increase the variance of the gradient estimator by only $\frac{2(m+n+1)}{mn+1}$ relative to dense Gaussian ES, or $0.098\%$ for a $4096\times4096$ matrix. We then introduce LOO-ROLL, a leave-one-out estimator that preserves the finite-rank population field while replacing EGGROLL's two antithetic evaluations per direction by one. At equal evaluation cost, LOO-ROLL halves estimator MSE in transformer blocks. At matched wall time across ten post-training settings and models up to 8B parameters, LOO-ROLL improves seven outcomes in individual paired tests, with no significant loss. On the GSM8K test set, accuracy increases from $38.1\%$ to $63.0\%$ at 0.6B and from $65.9\%$ to $80.0\%$ at 8B. Transformer measurements recover the predicted finite-rank variance, while the rank comparisons show no reproducible reward-based advantage for rank eight.
1 Introduction
The paper analyzes why low-rank EGGROLL can make large-scale ES practical despite finite-rank geometric and score mismatches, then introduces LOO-ROLL to improve evaluation efficiency. Its theory characterizes the population field and sampling error, while experiments show matched-time gains without significant losses.
- Motivation and method: EGGROLL replaces dense matrix perturbations with rank-r Gaussian products, reducing per-layer perturbation storage and arithmetic from O(mn) to O(r(m + n)).Batched computation lets population members share base weights and much of the inference machinery.
- Motivation and method: Rank-one perturbations occupy a zero-volume subset of matrix space despite identity covariance, and their dense-Gaussian score weighting creates a finite-rank score mismatch.These geometric and score issues are most pronounced at rank one.
- The finite-rank mean field: Finite-rank EGGROLL follows a resolvent-filtered gradient, can become nonconservative or destabilize an optimum, but is exact on quadratic objectives.The finite-rank correction is controlled under smoothness and is O(σ^2/r).
- Finite-population accuracy: Rank-r perturbations increase the dense-Gaussian variance factor by 2(m + n + 1)/r, so the relative penalty rapidly vanishes for wide matrices, even at rank one.The paper contrasts this O(m + n) contribution with the dense-Gaussian O(mn) contribution.
- LOO-ROLL: LOO-ROLL uses one fitness evaluation per direction and preserves EGGROLL’s finite-rank population field before score standardization.It replaces antithetic pairs with leave-one-out baselines.
- Numerical and transformer experiments: LOO-ROLL improves seven of ten matched-wall-time post-training outcomes, with no significant loss, and five gains survive Holm correction.The study spans ten settings, three model sizes, and two model families.
2 Related work
Related work situates EGGROLL among zeroth-order optimization, variance-reduction, structured-perturbation, and adaptive ES methods. The paper distinguishes its analysis of the coupled Gaussian-product matrix law and finite-rank population field from prior guarantees and methods.
- Zeroth-order and structured methods: EGGROLL differs from dense or coordinatewise zeroth-order methods by using batched low-rank Gaussian-product matrix perturbations.This construction produces the finite-rank population field analyzed in the paper.
- Variance reduction: LOO-ROLL specializes leave-one-out baselines to EGGROLL’s Gaussian-product population, extending variance-reduction ideas used in score-function and policy-gradient estimators.The related methods use baselines without changing estimator expectation.
- EGGROLL theory: Prior EGGROLL theory covered high-dimensional vanishing-radius limits and convergence to dense Gaussian ES at O(1/r), whereas this paper studies finite dimension, radius, and population.The paper frames its results as complementary to those asymptotic guarantees.
- Mathematical foundations: The paper’s operator identity preserves coordinate coupling in matrix products and yields a matrix differential operator whose resolvent connects to finite-rank optimization dynamics.This differs from scalar product-normal Stein identities that do not retain the matrix dependence.
- Structured perturbations: Unlike approaches that add isotropic Gaussian directions, the paper shows that identity covariance does not make EGGROLL spherically symmetric in vectorized matrix space.Its conservativity analysis isolates spherical symmetry as the stronger distributional requirement.
3 Setup
The setup defines EGGROLL’s finite-rank population field as the expected score-weighted update under Gaussian-product perturbations and compares it with the gradient of the smoothed objective. Because the low-rank law lacks the dense Gaussian score identity, the two fields need not coincide.
- Population field: The population field is the infinite-population update obtained by averaging fitness-weighted perturbations over the Gaussian-product law.A finite ES population estimates this field.
- Population field: Dense Gaussian ES uses Stein’s identity with score −E∞, but finite-rank Gaussian-product perturbations do not generally satisfy that identity.For rank below min(m, n), the law is lower-dimensional and lacks a full-dimensional density; at finite rank, −Er is not its exact score.
- Research questions: The setup asks whether the finite-rank population field equals the gradient of the finite-rank-smoothed objective or any other scalar objective.This motivates the later analysis of conservativity.
- Antithetic estimator: Antithetic sampling evaluates +Er and −Er for each of N directions, producing an unbiased estimator while canceling shared fitness components and even Taylor terms.The estimator uses two evaluations per direction.
- Analytical convention: The main analysis studies the raw-score field without the centering and standardization used in EGGROLL experiments; standardized updates are treated separately.This keeps the theoretical calculation explicit.
- Smoothed objective: The smoothed objective is a convolution with the perturbation law, while the population field adds the perturbation-weight factor and can be analyzed through Fourier modes.This calculation leads to the finite-rank transformation characterized later.
4 Characterization of the population field
Finite-rank EGGROLL first smooths the objective under Gaussian-product perturbations, then applies a resolvent that transforms each Fourier-mode direction anisotropically. This score mismatch can tilt modes away from the smoothed gradient, while the transformation is attenuating, controlled, and disappears in the dense-Gaussian limit.
- Finite-rank smoothing distinguishes Fourier modes with equal Frobenius norm but different singular-value profiles, unlike dense Gaussian smoothing.The product law retains separate singular-value dependence and is not invariant under arbitrary orthogonal transformations of the ambient matrix coordinates.
- At each frequency, the population field replaces the smoothed-gradient direction T with J(T) = (I_m + σ^2TT^⊤/r)^-1T, which can tilt the mode.The scalar characteristic-function factor is shared by both fields; the directional replacement produces the discrepancy.
- The population field equals the gradient of the finite-rank-smoothed objective transformed by the resolvent (I + σ^2L/r)^-1.The resolvent acts explicitly on each matrix frequency and separates smoothing from the additional effect of the mismatched dense-Gaussian score.
- The resolvent is an anisotropic low-pass filter: attenuation reaches one half when a singular value equals √r/σ, and unequal attenuation can change mode direction.The transformed singular coefficient increases up to this threshold and decreases beyond it.
- The population field remains globally nonnegatively aligned with the smoothed gradient, while score-mismatch error shrinks as σ^2/r decreases.Increasing rank or decreasing perturbation radius brings the population field closer to the smoothed-gradient field under the stated regularity conditions.
5 Conservativity of the population field
Finite-rank EGGROLL need not follow a conservative gradient field because its perturbation law lacks spherical symmetry. This nonconservative resolvent filtering can alter local optimization dynamics, including making a strict smoothed-objective maximum unstable under rank one.
- Universal conservativity holds exactly when the centered perturbation law is spherically symmetric.The characterization applies to all trigonometric-polynomial objectives in dimensions d ≥ 2.
- Finite-rank EGGROLL is not spherically symmetric for matrix dimensions m,n ≥ 2, so some objectives have nonconservative population fields.Its law is invariant under left-right orthogonal matrix transformations but not arbitrary rotations of the vectorized matrix coordinates.
- For the bounded cosine example, the population-field Jacobian is nonsymmetric at finite rank and nonzero radius, demonstrating failure of conservativity.The explicit m=n=2 construction uses T=diag(1,2).
- At r=σ=1, rank-one EGGROLL makes a strict local maximum of both the objective and its smoothed version linearly unstable.The rank-one population Jacobian has a positive eigenvalue, whereas sufficiently small ascent steps on the smoothed objective contract toward the maximum.
- In the local dynamics visualization, rank-one EGGROLL repels along one direction, while rank two and Richardson extrapolation are locally attracting.Positive and negative Jacobian real parts correspond to local repulsion and attraction, respectively.
6 Finite-rank dynamics near quadratic objectives
Finite-rank EGGROLL is exact for every quadratic objective, while smooth nonquadratic objectives incur a rank-dependent correction that can be controlled under local smoothness. The expansion also clarifies that its validity depends on perturbations remaining within a region where low-order Taylor approximations are accurate.
- Exactness on quadratic objectives: EGGROLL exactly recovers the true gradient for every quadratic objective at every rank r ≥ 1 and perturbation radius σ > 0.This includes arbitrary linear coefficients and cross-coordinate Hessian terms.
- Stability and approximation: The stability reversal arises when sampled perturbations reach regions where the objective is not well approximated by its quadratic Taylor polynomial, despite matching local curvature at the origin.The cosine example and its quadratic approximation have identical value, gradient, and Hessian at zero but different EGGROLL dynamics.
- Local expansion: The finite-rank correction for smooth objectives is proportional to σ^2/r, depends on third derivatives, vanishes on quadratics, and need not be a gradient field.Dense Gaussian smoothing contributes a separate leading correction shared with dense Gaussian ES.
- Scope of the approximation: The local expansion is accurate only when sampled perturbations stay in a region where the low-order Taylor approximation remains valid, which may fail in high dimension or with large higher-order derivatives.A numerically small σ may still produce a large typical displacement.
- Field-error bounds: Standard smoothness assumptions provide nonasymptotic field-error bounds, using Lipschitz gradients or sharper bounds when the Hessian is Lipschitz.The bounds retain the exact fourth moment of the Gaussian-product perturbation.
7 Finite-population variance in the local linear model
In the local affine model, finite-rank perturbations add a small variance correction to dense Gaussian ES despite their lower-dimensional support. For wide matrices, rank-one EGGROLL therefore estimates its own population field with nearly the same sampling error as dense Gaussian ES.
- Local linear model: Under the affine local model, each rank-r direction gives an unbiased estimator of the local gradient because the perturbation has identity covariance.The model removes curvature and stochastic fitness-evaluation noise to isolate directional sampling variability.
- Scope: The calculation conditions on stochastic fitness-evaluation randomness, isolating variability caused only by sampled perturbation directions.This is the scope of the affine variance analysis.
- Exact variance: The finite-rank variance factor is mn + 1 + 2(m + n + 1)/r, so the additional variance rapidly becomes negligible as matrix width grows.Dense Gaussian and Gaussian-product directions share the baseline factor mn + 1.
- Variance comparison: 0.098% is the relative variance increase of rank-one EGGROLL over dense Gaussian ES for a 4096 × 4096 matrix.The exact relative increase is 2(m + n + 1)/(mn + 1).
- Interpretation: The near-equality in sampling error does not imply equal population fields, because finite rank can change the mean field at nonzero perturbation radius.The variance comparison concerns a particular weighted fourth-moment sum, whereas singularity concerns perturbation support.
8 Approach to dense Gaussian ES at fixed radius
At fixed perturbation radius, increasing rank makes the EGGROLL field approach the dense-Gaussian ES field with an explicit 1/r expansion. The leading finite-rank error combines altered smoothing with score mismatch and generally cannot be improved beyond order 1/r.
- Geometric structure: The first finite-rank correction depends on singular-direction anisotropy through tr((TT ⊤)^2), whereas the dense-Gaussian leading term depends only on ∥T∥F.Thus anisotropy enters the fixed-radius error even when the leading dense-Gaussian term is isotropic in this representation.
- Uniform control: The fixed-radius expansion is established in a polynomially weighted uniform norm, allowing objectives with polynomial growth while scaling the error by the same growth rate.For bounded objectives, this reduces to an ordinary uniform bound.
- Rank expansion: The dense-Gaussian field is the zeroth-order term, the leading finite-rank correction is B_σf/r, and the remaining error is O(r^-2).When B_σf is nonzero, the 1/r rate cannot be improved in the weighted norm.
- Sources of error: The leading 1/r correction combines finite-rank versus dense-Gaussian smoothing with the score mismatch created by using the dense-Gaussian score.The remaining discrepancy is smaller by another factor of 1/r.
- Interpretation: The comparison isolates finite-rank effects because both fields use the same fixed perturbation radius and share the dense-Gaussian smoothing baseline.The resulting bounds can be combined with finite-population variance to obtain finite-iteration guarantees.
9 EGGROLL leads to LOO-ROLL
LOO-ROLL replaces EGGROLL’s antithetic pair evaluations with one evaluation per independent direction while preserving the finite-rank population field. At equal evaluation cost, its affine-model variance approaches half that of EGGROLL, motivating tests in transformer blocks and post-training.
- Computational trade-off: The practical motivation is that antithetic EGGROLL spends two evaluations per independent direction, whereas LOO-ROLL can spend the saved evaluation on another direction or optimization step.This is the stated fixed-budget and fixed-wall-time trade-off.
- Baseline mechanism: LOO-ROLL forms a leave-one-out baseline from the other population members, exploiting baseline subtraction without changing the estimator’s expectation.The baseline can be computed from the population mean already used for score centering.
- Estimator construction: LOO-ROLL preserves the finite-rank population field exactly before score standardization while using one forward evaluation per independent direction instead of two.Its leave-one-out baseline depends only on other directions and requires no additional fitness evaluations.
- Variance advantage: At equal cost of 2N fitness evaluations, LOO-ROLL uses 2N independent directions instead of N antithetic directions, and its affine-model variance approaches half of EGGROLL’s as N grows.The leave-one-out baseline adds only an order-N^-2 term to the order-N^-1 directional variance.
- Implementation: The method retains EGGROLL’s low-rank perturbations, shared batched forward computations, seed reconstruction, and weight-synchronization procedure.Thus the estimator change preserves the computational structure of low-rank ES.
10 Experiments: from finite-rank theory to LLM post-training
Experiments validate the finite-rank theory in controlled and transformer settings, then show that LOO-ROLL improves equal-time post-training outcomes. Rank increases reduce variance or improve per-update progress, but their computational cost usually removes the benefit.
- 10.1.1 Controlled tests of the theoretical predictions: Controlled tests reproduce the predicted finite-rank dynamics: rank one changes the counterexample’s largest Jacobian real part from −0.00444 at rank two to 0.02925 at rank one, while sampling variance agrees within 0.9%.The eigenvalue change corresponds to local stability becoming repulsion, and Monte Carlo trajectories also match the predicted population-noise transition.
- 10.1.2 Is the rank-one surcharge visible in transformer blocks?: Transformer-block MSE ratios follow affine predictions at ranks 1, 2, 4, and 8, while mean-gradient cosine similarity remains 0.980–0.984 across ranks and radii.For a 16 × 16 attention block, observed ratios are 1.264, 1.127, 1.089, and 1.028 versus predictions 1.257, 1.128, 1.064, and 1.032; the predicted surcharge falls below 0.1% at width 4096.
- 10.2.1 Does higher rank improve training at matched cost?: Rank-eight updates achieve lower next-token loss at matched iterations across four model sizes, but the advantage disappears or reverses at matched wall time except at 1.7B.Rank eight retains an equal-time advantage only at 1.7B because its additional iteration cost usually outweighs per-update gains.
- 10.3 LOO-ROLL halves evaluation cost and improves equal-time performance: At equal evaluation cost, doubling LOO-ROLL directions approximately halves update MSE, while equal-direction MSE is at most 5.2% higher and reaches 0.8% at N = 128.The equal-evaluation comparison uses twice as many LOO-ROLL directions as antithetic EGGROLL.
- 10.3 LOO-ROLL halves evaluation cost and improves equal-time performance: LOO-ROLL improves equal-wall-time post-training outcomes, with seven individual paired-test gains, five surviving Holm correction, and no significant loss across ten settings.The comparison uses nearly equal runtimes, differing from EGGROLL by -0.3% to +2.7%.
- 10.3 LOO-ROLL halves evaluation cost and improves equal-time performance: GSM8K accuracy rises from 38.10% to 63.00% at 0.6B and from 65.87% to 80.00% at 8B with LOO-ROLL at equal wall time.The gains are positive for every training seed at all three Qwen3 sizes, with no setting significantly favoring antithetic EGGROLL.
11 Conclusion
Finite-rank EGGROLL changes both mean dynamics and sampling behavior: its joint network field is resolvent-filtered, while low-rank variance remains structured and analyzable. The experiments connect this analysis to LLM post-training and motivate LOO-ROLL as a more evaluation-efficient estimator.
- 11 Conclusion: Finite-rank EGGROLL addresses both mean-field direction and finite-population accuracy, explaining how rank one can retain nearly dense-ES sampling efficiency despite differing mean dynamics.The variance contribution grows with m + n, versus the dense-Gaussian contribution of order mn.
- 11 Conclusion: LOO-ROLL halves estimator MSE at equal evaluation cost and improves seven of ten matched-wall-time settings, with no significant loss.On GSM8K, accuracy rises from 38.1% to 63.0% at 0.6B and from 65.9% to 80.0% at 8B.
- 11 Conclusion: The network-level results extend field-error and variance statements across independently perturbed parameter blocks under smoothness assumptions.The joint treatment concatenates blocks while retaining within-block variance terms.
- 11 Conclusion: For coupled network parameters, independent block perturbations yield a block-diagonal resolvent Jr,σ = diagq(I + σ^2Lq/rq)^−1 applied to the smoothed gradient.The operator is self-adjoint, positive, contractive, and firmly nonexpansive on the direct-sum space.
B Secondary corrections suggested by the analysis
The analysis suggests rank extrapolation and control-variate corrections, but both secondary estimators introduce costs or variance that prevent practical gains in the reported studies. Standard smoothness results instead provide convergence guarantees for the original or smoothed objectives.
- B.1 Rank extrapolation: Extrapolation cancels a leading mean-field term but requires estimating two fields, while its weights amplify sampling errors.This trade-off motivates treating the method as a theoretical correction rather than a practical contribution.
- B Secondary corrections suggested by the analysis: At a fixed budget of 256 fitness evaluations, rank extrapolation produces 7.63–8.67 and 2.35–2.62 times ordinary rank-one EGGROLL’s MSE for ranks (1,2) and (1,4).The leading bias is below sampling resolution, so cancellation does not repay splitting the population between ranks.
- B Secondary corrections suggested by the analysis: The control-variate estimator is unbiased for the finite-rank field, but tests on Qwen3-0.6B and Qwen3-1.7B found all four means favoring ordinary EGGROLL.The results do not establish a practical benefit from maintaining the reference estimate and its additional evaluations.
- B Secondary corrections suggested by the analysis: Standard smoothness theorems separate optimization progress, mean-field error, and finite-population noise under step-size and population-size conditions.The supplied results include explicit finite-population complexity and convergence statements relative to both the original loss and dense Gaussian ES.
C.1 Finite-population convergence
Finite-population convergence is governed by two errors: sampling fluctuations around the population field and the field’s discrepancy from the chosen gradient reference. Smoothness-based bounds quantify both effects and distinguish comparison with the original loss from comparison with dense Gaussian ES.
- C.1 Finite-population convergence: Convergence guarantees require controlling finite-population noise, mean-field error, and optimization progress under a step-size constraint involving relative population noise.The analysis separates multiplicative noise from additive evaluation noise that can persist when the mean update vanishes.
- C.1 Finite-population convergence: Relative to the original loss, field error combines smoothing and finite-rank score mismatch; relative to dense Gaussian ES at the same radius, shared smoothing cancels.The latter comparison leaves a mean-field difference of O(1/r).
- C.1 Finite-population convergence: Finite-rank EGGROLL’s mean-field contribution to squared-stationarity is O(r^−2) relative to dense Gaussian ES when the leading coefficient is nonzero.For bounded fitness, no trajectory-moment condition is needed.
- C.1 Finite-population convergence: The rank-one fourth moment consists of the three dense-Gaussian pairings plus six factorization-specific terms scaled by 1/r.This fourth-moment structure supplies the rank-dependent variance and field-error corrections.
- C.1 Finite-population convergence: Transformer-block measurements show rank-one MSE 26–32% above dense ES, falling to 3–4% above it at rank eight, while cosine similarity remains near 0.98.These results make finite-rank variance visible while showing little change in the mean direction over the tested radii.
E.3 Controlled numerical experiments
Controlled experiments reproduce the paper’s exact variance, resolvent, stability, and rank-expansion predictions. Finite populations can expose rank-dependent instability at small sample sizes, while larger populations and wider matrices reduce the relative rank-one contribution.
- E.3 Controlled numerical experiments: Direct numerical calculations reproduce the covariance tensor, resolvent multiplier, Hessian results, cubic identity, and 1/r rank expansion.These checks verify the algebra behind the theoretical predictions before finite-sampling validation.
- E.3 Controlled numerical experiments: The empirical single-direction variance is within 0.9% of the exact formula at every rank, and the largest Jacobian real part shifts from 0.02925 at rank one to −0.00444 at rank two.The measurements recover the predicted stability transition.
- E.3 Controlled numerical experiments: At N = 16, Monte Carlo trajectories match the exact mean-square prediction, with rank one unstable while rank eight and dense Gaussian ES contract.The figure validates both one-step ratios and full trajectories.
- E.3 Controlled numerical experiments: With few directions, the κr/N term can cross the stability boundary; increasing N removes the difference, and increasing matrix width makes the rank-one contribution small relative to mn + 1.The experiment uses 8 × 8 dynamics with one-step trials and 25-step trajectories.
F Effect of population centering and standardization
Population centering preserves each realized update direction, while common standardization only rescales it. Same-population normalization can nevertheless bias the expected field, although the measured directional effect is small at training population size.
- Population centering and common standardization: Population centering cancels in every antithetic pair, and dividing all scores by one positive common scale changes only the update magnitude.The resulting update is exactly a scalar multiple of the unstandardized update.
- Expected-field bias: A same-population normalizer can change the expected update field when its inverse scale is statistically coupled with estimator noise.Independence, or conditional mean-zero noise given the scale information, removes this bias.
- Finite-population effect: Same-population standardization changes the raw population field by an O(N^-1) term apart from the scalar factor E[α_N].The bound follows under ordinary averaging and empirical-scale concentration conditions.
- Transformer-block audit: At N = 128, rank-one raw and standardized means have cosine similarity 0.99970, while the dense control reaches 0.99990.The audit averages two transformer blocks, four radii, and three seeds; higher cosine and lower residual indicate smaller directional effects.
- Transformer-block audit: At N = 128, the rank-one standardized field retains a 2.42% directional residual, compared with 1.44% for the dense control.The residual is the standardized component unexplained by rescaling the raw field.