Source-linked AI summary

EZmocks: extending the Zel'dovich approximation to generate mock galaxy catalogues with accurate clustering statistics

Chia-Hsun Chuang, Francisco-Shu Kitaura, Francisco Prada, Cheng Zhao, Gustavo Yepes

arXiv:1409.1124v2astro-ph.CO

TL;DR

Large-scale galaxy surveys need accurate mock catalogues, but full N-body simulations are computationally expensive and simpler log-normal models miss higher-order structure. EZmocks extend the Zel’dovich approximation with effective prescriptions for bias and nonlinear effects, achieving close agreement with N-body clustering statistics while remaining fast and practical for massive catalogue production.

  • Problem

    Large surveys require accurate, scalable mock catalogues, while full N-body simulations are costly and log-normal mocks do not capture higher-order structure.

  • Method

    EZmocks extend the Zel’dovich approximation with PDF mapping, density thresholding and saturation, BAO and small-scale power corrections, and effective biasing prescriptions.

  • Results

    EZmocks reproduce one-, two-, and three-point clustering statistics nearly indistinguishably from full N-body solutions, with one 960^3-grid mock taking about 5 minutes on 16 cores.

  • Takeaways & Limitations

    EZmocks provide a reliable and practical way to generate massive catalogues for large-scale-structure analysis, covariance estimation, and future survey forecasts.

  • Takeaways & Limitations

    EZmocks fit clustering statistics through explicit prescriptions rather than explicit physical bias relations, and remain complementary to indispensable high-quality N-body reference catalogues.

Abstract

from arXiv · show

We present a new methodology to generate mock halo or galaxy catalogues, which have accurate clustering properties, nearly indistinguishable from full $N$-body solutions, in terms of the one-point, two-point, and three-point statistics. In particular, the agreement is remarkable, within $1\%$ up to $k=0.55$ $h$Mpc$^{-1}$ and down to $r=10$ $h^{-1}$Mpc, for the power spectrum and two-point correlation function respectively, while the bispectrum agrees in general within $20\%$ for different scales and shapes. Our approach is based on the Zel'dovich approximation, however, effectively including with the simple prescriptions the missing physical ingredients, and stochastic scale-dependent, non-local and nonlinear biasing contributions. The computing time and memory required to produce one mock is similar to that using the log-normal model. With high accuracy and efficiency, the effective Zel'dovich approximation mocks (EZmocks) provide a reliable and practical method to produce massive mock galaxy catalogues for the analysis of large-scale structure measurements.

1 INTRODUCTION

Growing galaxy surveys require massive mock catalogues that preserve clustering accurately without the prohibitive computational demands of full N-body simulations. EZmocks extend the Zel’dovich approximation with effective biasing and nonlinear corrections to provide efficient catalogues spanning one-, two-, and three-point statistics.

  • Motivation: Mock catalogues are essential for analysing survey clustering, but full N-body simulations are often impractical because their runtime and memory demands are prohibitive.Upcoming surveys require precise mocks across large parameter spaces, motivating simpler and more efficient methods.
  • Existing methods: Log-normal mocks reproduce two-point statistics but neglect matter displacements and cosmic-web structure, producing large deviations in higher-order statistics.
  • Existing methods: Perturbation-theory and statistical halo methods reduce costs, but approximate gravity limits their accuracy in nonlinear regimes and halo formation.
  • Contribution: EZmocks extend the Zel’dovich approximation with stochastic, scale-dependent, non-local, and nonlinear biasing contributions to reproduce one-, two-, and three-point clustering statistics.
  • Contribution: EZmock generation requires three FFTs for displacement fields, nearby halo population, and only a few grid-sized arrays for memory.

2 REFERENCE SIMULATIONS AND COMPUTATIONAL REQUIREMENTS

The study tests EZmocks against a BigMultiDark reference halo catalogue and measures the computational resources needed to generate them. A 960^3 grid mock can be produced in under five minutes on one 16-core node.

  • Reference simulations: The reference catalogue comes from a BigMultiDark simulation at z = 0.5618, using 3840^3 particles in a (2500 h^-1 Mpc)^3 volume with Planck ΛCDM parameters.
  • Computational requirements: A 960^3-grid EZmock takes less than five minutes to construct on one 16-core, 64 GB node, using less than 32 GB of memory.The Curie system used shared-memory multiprocessing to accelerate computation.

3 METHODOLOGY

EZmocks extend the Zel’dovich approximation with explicit prescriptions for PDF mapping, stochastic and nonlinear bias, scale-dependent power-spectrum corrections, BAO enhancement, and velocities. The methodology converts the ZA density field into halo catalogues through rank ordering, scatter, thresholding, saturation, spectral modifications, and mass assignment.

  • 3 METHODOLOGY: EZmocks augment the Zel’dovich approximation with stochastic, nonlinear, non-local, and scale-dependent biasing contributions to model halo catalogues.The approach assumes ZA captures the large-scale cosmic-web structure while subsequent prescriptions supply missing physical ingredients.
  • 3 METHODOLOGY: The pipeline generates a ZA density field, maps the BigMD halo PDF by rank ordering, adds scatter, fits amplitudes and shapes, enhances BAO, and computes velocities.These steps are recursively applied until convergence, combining PDF mapping, thresholding, saturation, spectral tilting, BAO enhancement, and velocity assignment.
  • 3.1 Generation of the Zel’dovich density field: The ZA displacement field maps Lagrangian positions to Eulerian positions, using Fourier-space density perturbations to construct the field at z = 0.5618.The chosen 960^3 grid is near optimal: smaller grids damp the BAO peak, while larger grids produce similar results at higher computational cost.
  • 3.2 PDF mapping scheme: Rank ordering transfers Poisson-sampled BigMD halo counts onto the ZA grid, then populates neighboring cells through a CIC distribution.The CIC assignment restricts each halo to the eight neighboring cells with probabilities determined by relative distances.
  • 3.3 Adding scatter to the PDF mapping: Scatter models stochastic tracer uncertainty, while density thresholding and saturation adjust three-point statistics and power-spectrum amplitude.The scatter uses a Gaussian variable and an exponential branch for negative values to avoid negative densities; λ is fixed to 10 without affecting performance.
  • 3.5 Nonlinear effects correction: Residual nonlinear effects are corrected by enhancing BAO and small-scale power and by modifying the input spectrum’s shape with scale-dependent adjustments.The correction addresses small-scale power loss and BAO weakening caused by approximate evolution and halo-position uncertainties of a few Mpc.
  • 3.6 Peculiar velocity: Peculiar velocities combine coherent ZA displacement-based motion with a Gaussian random dispersion term.The coherent component is proportional to the displacement field, with B representing linear growth and λ′ the Gaussian width.
  • 3.7 Mass assignment: Halo masses are assigned by probabilistically dividing the BigMD catalogue into multiple mass-based sub-catalogues rather than using sharp mass cuts.The sub-catalogues are designed to contain comparable numbers of haloes and are used to construct mass-inclusive EZmocks.

4 VALIDATION OF THE METHOD

EZmock validation compares one-, two-, and three-point statistics, mass dependence, and the halo PDF against BigMD. The method reproduces real-space clustering accurately, while redshift-space velocities and small-scale three-point statistics remain less precise.

  • 4.1 Probability distribution function: The halo PDF is sufficiently matched to reproduce higher-order statistics, although its high-occupancy tail is difficult to fit because of small-number statistics.The PDF mapping fixes the number density by construction, but CIC halo placement means the reference PDF is not guaranteed to be exactly restored.
  • 4.2 Power spectrum and two-point correlation function: Within 1%, EZmock matches the BigMD real-space power spectrum up to k = 0.55 h Mpc^-1 and the two-point correlation function down to 10 h^-1 Mpc.The power-spectrum comparison excludes scales above k = 0.6 h Mpc^-1 because of the 960^3 grid resolution.
  • 4.2 Power spectrum and two-point correlation function: In redshift space, the simple linear-plus-Gaussian velocity model fits worse than in real space but recovers the Kaiser factor well enough for most practical applications.The result indicates remaining room for improvement in the velocity assignment model.
  • 4.3 Bispectrum: The bispectrum is generally reproduced within 20% across the tested real- and redshift-space configurations, with better coverage in redshift space.In real space, the stated accuracy applies to configurations with k2 = 2k1 = 0.1 and 0.2 h Mpc^-1; all tested configurations meet it in redshift space.
  • 4.4 Mass vs. bias: EZmock restores the mass-dependent power-spectrum trend across 10 number densities and reconstructs the mass function by construction.The mass-bias dependence uses five mass bins, and greater accuracy may be obtained with more bins.

5 CONCLUSION AND DISCUSSION

EZmocks extend the Zel’dovich approximation with effective models to generate efficient mock catalogues that reproduce one-, two-, and three-point clustering statistics. They are complementary to full-gravity catalogues and improve substantially over log-normal mocks at comparable computational cost.

  • EZmocks provide accurate one-, two-, and three-point clustering statistics while generating massive catalogues efficiently.A 960^3-grid mock takes about five minutes on 16 cores.
  • Effective models compensate for missing physical contributions and absorb nonlinear growth and deterministic and stochastic halo-bias effects.The implementation combines PDF mapping with scattering, density threshold and saturation, and enhancements to BAO and small-scale power.
  • The method remains complementary to full-gravity simulations, which are still needed to produce accurate reference catalogues.EZmocks can then reproduce those reference catalogues in large numbers.
  • EZmock improves significantly on log-normal mocks at comparable computational cost and is practical for covariance estimation and future survey forecasts.The paper positions it as a practical method for producing many catalogues for large-scale-structure analyses.
Loading 1409.1124v2…