Source-linked AI summary
Learning to Predict the Cosmological Structure Formation
Siyu He, Yin Li, Yu Feng, Shirley Ho, Siamak Ravanbakhsh, Wei Chen, Barnabás Póczos
TL;DR
Accurately predicting cosmological structure requires simulations, but N-body methods are computationally expensive. The paper develops D3M to predict nonlinear structure from linear perturbation theory and finds it outperforms 2LPT while generalizing beyond its training cosmologies.
Problem
Predicting Universe-scale structure requires accurate simulations, but N-body simulation is computationally expensive.
Method
D3M is a deep neural network that maps the ZA displacement field to FastPM-like structure-formation predictions.
Results
D3M outperforms 2LPT in point-wise, 2-point, and 3-point comparisons and generalizes to significantly different cosmological parameters.
Takeaways & Limitations
Deep learning provides a practical and accurate alternative to traditional approximate simulations of cosmological structure formation.
Takeaways & Limitations
Replacing FastPM with exact N-body or higher-resolution simulations is expected to improve D3M performance.
Abstract
from arXiv · showhide
Matter evolved under influence of gravity from minuscule density fluctuations. Non-perturbative structure formed hierarchically over all scales, and developed non-Gaussian features in the Universe, known as the Cosmic Web. To fully understand the structure formation of the Universe is one of the holy grails of modern astrophysics. Astrophysicists survey large volumes of the Universe and employ a large ensemble of computer simulations to compare with the observed data in order to extract the full information of our own Universe. However, to evolve trillions of galaxies over billions of years even with the simplest physics is a daunting task. We build a deep neural network, the Deep Density Displacement Model (hereafter D$^3$M), to predict the non-linear structure formation of the Universe from simple linear perturbation theory. Our extensive analysis, demonstrates that D$^3$M outperforms the second order perturbation theory (hereafter 2LPT), the commonly used fast approximate simulation method, in point-wise comparison, 2-point correlation, and 3-point correlation. We also show that D$^3$M is able to accurately extrapolate far beyond its training data, and predict structure formation for significantly different cosmological parameters. Our study proves, for the first time, that deep learning is a practical and accurate alternative to approximate simulations of the gravitational structure formation of the Universe.
Significance Statement
D3M is presented as a deep neural network for predicting cosmological structure formation, offering an accurate and computationally efficient alternative to expensive simulations. It generalizes beyond its training cosmological parameters.
- D3M predicts Universe structure formation and is presented as an accurate alternative to traditional approximate simulations.The model is intended to provide a faster approach to prediction than computationally expensive N-body simulations.
- The study applies deep learning to generate complex 3D simulations in cosmology.
- D3M accurately extrapolates beyond its training data and generalizes to significantly different cosmological parameters.The authors report that training on a single cosmological parameter set can support predictions for new parameter sets.
Setup
D3M transforms first-order displacement information into approximate nonlinear structure formation, using ZA inputs and FastPM targets. The setup compares this approach with 2LPT and trains on 10,000 simulation pairs.
- D3M takes the displacement field from ZA as input and learns to predict the target displacement field produced by FastPM.ZA evolves particles along linear trajectories, while FastPM provides an accurate approximate N-body target.
- 2LPT is selected as the reference model because it efficiently generates relatively accurate large-scale structure descriptions.It bends particle trajectories using a quadratic correction and is commonly used for cosmological simulation ensembles.
- 10,000 pairs of ZA inputs and FastPM targets are generated from simulations of 32^3 particles in a 128 h^-1Mpc volume.The mean particle separation is 4 h^-1Mpc per dimension.
- Training on displacement fields avoids ambiguity that arises because different nonlinear displacement fields can produce identical density fields.The authors report that using density fields as both input and target produced less comparable results.
Results and Analysis
D3M produces displacement fields close to FastPM ground truth and identifies cosmic-web structures such as clusters, filaments, and voids. Its average relative error is substantially lower than 2LPT's.
- D3M predictions reveal cosmic-web structures including clusters, filaments, and voids in the resulting point-cloud representation.
- 2.8% average relative error for D3M versus 9.3% for 2LPT in displacement-field predictions.Across 1,000 simulations, the maximum relative errors are 1.10 for D3M and 4.23 for 2LPT.
- Denser regions have higher errors for all methods, consistent with greater nonlinearity in structure formation.
2-Point Correlation Comparison.
The paper evaluates D3M and 2LPT using two-point statistics of density and displacement fields against FastPM. D3M closely reproduces the target across large to semi-nonlinear scales and outperforms 2LPT, with deviations emerging mainly at the smallest scales.
- Scale dependence: Denser regions have higher prediction errors for D3M, 2LPT, and ZA, indicating that highly nonlinear regions are harder to predict.The point-wise comparison is illustrated using two-dimensional slices of particle distributions and displacement vectors.
- Metrics: Two-point performance is measured with transfer function T(k) and correlation coefficient r(k) against FastPM.T(k) captures amplitude discrepancies, while r(k) captures phase discrepancies; 1-r^2 measures stochasticity.
- 2LPT comparison: 2LPT’s displacement transfer function exceeds 1 near k ≈0.35 hMpc−1 because it overestimates displacement power at small scales.A sharp power drop near the voxel scale results from smoothing that erases smaller-scale power.
- D3M results: D3M transfer functions differ from 1 by 0.4% for k ≲0.4 hMpc−1, increasing to 2% for density and 4% for displacement at k ≈0.7 hMpc−1.The reported stochasticity is approximately 10−3 and 10−2 for most scales, with correlation above 90% down to k = 0.7 hMpc−1.
- D3M results: D3M reproduces structure formation from large to semi-nonlinear scales and significantly outperforms 2LPT in the two-point analysis.D3M only begins to deviate from FastPM at fairly small scales, where deeply nonlinear evolution is difficult for current analytical theories.
3-Point Correlation Comparison.
The paper compares the three-point correlation functions of D3M and 2LPT with FastPM through multipole moments across triangle configurations. D3M is closer to FastPM, has smaller error bars, and yields a substantially lower relative residual than 2LPT.
- Comparison setup: The 3PCF comparison uses multipole moments ζℓ(r1, r2) across several triangle configurations, with results averaged over 10 test simulations.The analysis uses 10 radial bins with Δr = 5 h−1Mpc and reports standard-deviation error bars.
- 3PCF comparison: D3M’s 3PCF is closer to FastPM than 2LPT’s and has smaller error bars.The comparison uses ratios of predicted to target binned multipole coefficients.
- Quantitative result: 0.79% and 7.82% are the mean relative 3PCF residuals for D3M and 2LPT, respectively, compared with FastPM.The paper describes D3M’s 3PCF accuracy as an order of magnitude better than 2LPT’s.
- Interpretation: D3M’s improved 3PCF accuracy indicates that it captures non-Gaussian structure formation better than 2LPT.The 3PCF is used to assess correlations among three locations in configuration space.
Generalizing to New Cosmological Parameters
D3M trained on one cosmological-parameter set predicts structure formation for substantially different As and Ωm values. This suggests simulations across broader parameter ranges can be produced with limited additional training data.
- D3M trained on a single parameter set predicts structure formation for widely different As and Ωm choices.
- Lower As or Ωm produces larger distribution differences, while higher values produce more nonlinear displacements that make prediction harder.Figure 4 compares particle distributions and displacement fields against the training cosmology.
- D3M’s generalization suggests simulations spanning diverse cosmological parameters may require minimal training data.
- Changing As to 0.2A0 and 1.8A0 leaves D3M’s average relative displacement error below 4% per voxel.The same-parameter training and testing error is below 3%.
Varying Primordial Amplitude of Scalar Perturbations As.
When testing beyond the training cosmology, D3M generally provides more accurate structure predictions than 2LPT, with the clearest advantages for larger As. At the largest scales for smaller As, 2LPT can perform better.
- For As = 1.8A0, D3M performs much better than 2LPT in transfer-function and correlation-coefficient tests.
- For As = 0.2A0, 2LPT performs better at the largest scales, while D3M is more accurate at scales larger than k = 0.08 hMpc−1.
- D3M’s 3PCF predictions are notably better than 2LPT’s for larger As and only slightly better for smaller As.
Varying matter density parameter Ωm.
D3M is evaluated on Ωm values substantially different from training, using two-point statistics and comparisons with 2LPT. Its advantage is strongest for Ωm = 0.5 and at smaller scales for Ωm = 0.1.
- Table 1 summarizes the analysis of D3M and 2LPT across the tested cosmological settings.
- For Ωm = 0.5, D3M outperforms 2LPT at all scales in two-point density statistics.
- For Ωm = 0.1, D3M outperforms 2LPT on smaller scales where k > 0.1 hMpc−1.
- The mean relative 3PCF residuals for D3M are 1.7% at Ωm = 0.5 and 1.2% at Ωm = 0.1.
- The corresponding 2LPT mean relative 3PCF residuals are 7.6% at Ωm = 0.5 and 1.7% at Ωm = 0.1.
Conclusions
D3M accurately predicts FastPM large-scale structure across scales, improves on 2LPT in the nonlinear regime, and generalizes to cosmologies outside its training set. The method remains tied to FastPM simulations in the reported evaluation.
- D3M accurately predicts FastPM large-scale structure at all scales and models the nonlinear regime more accurately than 2LPT.
- D3M generalizes well to test simulations with As and Ωm values significantly different from those used for training.
- D3M learns a nonlinear mapping from first-order perturbation theory to FastPM beyond what higher-order perturbation theories currently achieve.
- Replacing FastPM with exact N-body simulations is expected to improve the method’s performance.
Materials and Methods
The study uses 10,000 ZA–FastPM simulation pairs and divides them into training, validation, and test sets. D3M is a three-dimensional U-Net trained with Adam, mean-squared error, and L2 regularization.
- Materials and Methods: 10,000 simulations provide ZA–FastPM input-output pairs covering an effective volume of 20 (Gpc/h)3.The volume is comparable to a large spectroscopic survey such as DESI or EUCLID.
- Materials and Methods: 80%, 10%, and 10% of the full simulation dataset are assigned to training, validation, and test sets, respectively.
- Materials and Methods: D3M uses a three-dimensional U-Net with 15 convolution or deconvolution layers and approximately 8.4 × 10^6 trainable parameters.The architecture generalizes the standard U-Net to three-dimensional data.
- Materials and Methods: Training uses Adam with a learning rate of 0.0001, mean-squared error loss, and L2 regularization with coefficient 0.0001.The Adam exponential decay rates are 0.9 and 0.999 for the first and second moments.
Details of the D3M Architecture.
D3M uses a three-dimensional convolutional architecture with periodic-boundary handling to predict displacement fields. Its loss connects integrated squared error with Fourier-space amplitude and phase similarity, while periodic padding improves performance and convergence.
- Details of the D3M Architecture.: The contracting path uses two convolutional blocks with stride-1 convolutions and stride-2 down-sampling, while feature channels double at each down-sampling step.The convolutions use 3×3×3 filters with periodic padding.
- Details of the D3M Architecture.: The expansive path halves feature channels before concatenating them with corresponding contracting-path features, then produces the final three-dimensional displacement field.The final conversion uses a 1×1×1 convolution; earlier convolutions use ReLU activation and batch normalization.
- Details of the D3M Architecture.: Periodic padding preserves the simulation box’s toroidal boundary connections that constant or reflective padding would disrupt.The periodic condition is expressed as x_i+L = x_i, with L denoting the box periodicity.
- Details of the D3M Architecture.: Periodic padding significantly improves performance and expedites convergence compared with constant padding in the same network.
- Details of the D3M Architecture.: D3M is trained to minimize mean-squared particle-displacement error, which is proportional to integrated squared error.Parseval’s theorem rewrites this loss in Fourier space.
- Details of the D3M Architecture.: The Fourier-space loss jointly captures transfer-function amplitude and correlation-coefficient phase similarity, approaching zero as both approach one.Here q is Lagrangian position, k its wavevector, T the transfer function, and r the correlation coefficient.
- Details of the D3M Architecture.: The implementation and training-data-generation code are publicly available in the cited repositories.