Source-linked AI summary
The CAMELS project: Cosmology and Astrophysics with MachinE Learning Simulations
Francisco Villaescusa-Navarro, Daniel Anglés-Alcázar, Shy Genel, David N. Spergel, Rachel S. Somerville, Romeel Dave, Annalisa Pillepich, Lars Hernquist, Dylan Nelson, Paul Torrey, Desika Narayanan, Yin Li, Oliver Philcox, Valentina La Torre, Ana Maria Delgado, Shirley Ho, Sultan Hassan, Blakesley Burkhart, Digvijay Wadekar, Nicholas Battaglia, Gabriella Contardo, Greg L. Bryan
TL;DR
CAMELS addresses the need for theory predictions that support extracting small-scale cosmological information while accounting for uncertain baryonic physics. It constructs a large simulation suite spanning cosmology and astrophysics, combines it with machine-learning applications, and demonstrates broad scientific potential alongside explicit scope and computational limitations.
Problem
Small-scale cosmological information is difficult to use because non-Gaussian fields lack a known optimal summary statistic and baryonic effects are poorly understood.
Method
CAMELS combines thousands of N-body and (magneto-)hydrodynamic simulations with varied cosmological and astrophysical parameters and applies machine-learning methods to their outputs.
Results
CAMELS supports applications including symbolic regression, GAN-based data generation, dimensionality reduction, and anomaly detection, while its simulations show substantial baryonic variation and differences between IllustrisTNG and SIMBA.
Takeaways & Limitations
The suite provides a proof-of-concept platform for developing methods that extract cosmological information while marginalizing over baryonic effects.
Takeaways & Limitations
CAMELS varies only two cosmological and four astrophysical parameters, with Ω_b fixed when Ω_m varies, and its applicability remains limited in some regimes.
Abstract
from arXiv · showhide
We present the Cosmology and Astrophysics with MachinE Learning Simulations --CAMELS-- project. CAMELS is a suite of 4,233 cosmological simulations of $(25~h^{-1}{\rm Mpc})^3$ volume each: 2,184 state-of-the-art (magneto-)hydrodynamic simulations run with the AREPO and GIZMO codes, employing the same baryonic subgrid physics as the IllustrisTNG and SIMBA simulations, and 2,049 N-body simulations. The goal of the CAMELS project is to provide theory predictions for different observables as a function of cosmology and astrophysics, and it is the largest suite of cosmological (magneto-)hydrodynamic simulations designed to train machine learning algorithms. CAMELS contains thousands of different cosmological and astrophysical models by way of varying $Ω_m$, $σ_8$, and four parameters controlling stellar and AGN feedback, following the evolution of more than 100 billion particles and fluid elements over a combined volume of $(400~h^{-1}{\rm Mpc})^3$. We describe the simulations in detail and characterize the large range of conditions represented in terms of the matter power spectrum, cosmic star formation rate density, galaxy stellar mass function, halo baryon fractions, and several galaxy scaling relations. We show that the IllustrisTNG and SIMBA suites produce roughly similar distributions of galaxy properties over the full parameter space but significantly different halo baryon fractions and baryonic effects on the matter power spectrum. This emphasizes the need for marginalizing over baryonic effects to extract the maximum amount of information from cosmological surveys. We illustrate the unique potential of CAMELS using several machine learning applications, including non-linear interpolation, parameter estimation, symbolic regression, data generation with Generative Adversarial Networks (GANs), dimensionality reduction, and anomaly detection.
3. SIMULATIONS
CAMELS combines thousands of N-body and (magneto-)hydrodynamic simulations spanning cosmological and astrophysical parameters to characterize observables and support machine-learning applications. Its suites reveal broad variation and important differences between IllustrisTNG and SIMBA, especially in baryonic effects on matter clustering and halo baryon content.
- Simulation suite: 4,233 simulations evolve dark matter and, for hydrodynamic runs, gas resolution elements in periodic (25 h^-1Mpc)^3 boxes.The suite includes approximately half N-body and half state-of-the-art (magneto-)hydrodynamic simulations using IllustrisTNG and SIMBA subgrid physics.
- Parameter coverage: CAMELS varies cosmology and feedback, including extreme AGN, supernova, and no-feedback models, to span diverse physical conditions.The extreme set fixes cosmology and initial seed while varying feedback strength, including AAGN1 = 100, ASN1 = 100, and zero-feedback cases.
- Matter clustering: 40%–70% matter-power-spectrum variation and up to ∼50% hydrodynamic-to-N-body suppression show that baryonic effects substantially alter small-scale clustering.The matter-spectrum range runs from 40% at k ≃30 hMpc^-1 to 70% at k ≃1 hMpc^-1; strong feedback can produce ∼50% effects at k ∼10 hMpc^-1.
- Sources of variation: Cosmology and astrophysics dominate small-scale variation, while cosmic variance dominates large-scale variation for many properties.For star-formation-rate density, cosmic variance contributes only a small fraction of the variation, while extreme models can produce high-redshift distribution tails.
- Matter clustering: ∼50%–130% variation in the gas power spectrum is larger than the total-matter variation, with SIMBA generally showing lower gas clustering and a broader percentile range.The reported difference is attributed to distinct feedback implementations and effective feedback strengths.
- Galaxy and halo properties: SIMBA and IllustrisTNG differ most in halo baryon fractions, with SIMBA below IllustrisTNG across the full reduced-halo-mass range.Both suites nevertheless show low fractions in low-mass halos, a peak near ∼10^12 h^-1M⊙, and declining fractions at higher masses.
4.2. Variations due to cosmology and astrophysics
CAMELS examines how cosmological and astrophysical parameters alter the matter power spectrum and star-formation rate density, finding strong model-dependent responses, especially between IllustrisTNG and SIMBA. It then motivates machine-learning emulation of these nonlinear dependencies.
- Matter power spectrum: IllustrisTNG and SIMBA agree well on large-scale cosmological responses but differ on smaller-scale matter-power-spectrum responses.These differences persist even when varying cosmological parameters.
- Matter power spectrum: All astrophysical parameters affect the matter-power-spectrum amplitude and shape, with especially large IllustrisTNG–SIMBA differences for ASN2.Decreasing ASN2 raises small-scale power in IllustrisTNG but lowers it in SIMBA; SIMBA also shows percent-level large-scale effects from wind and jet velocities.
- Matter power spectrum: Cosmological parameter changes produce larger matter-power-spectrum variations than astrophysical changes, partly because the explored cosmological ranges are much broader.The matter power spectrum is dominated by dark-matter clustering, which responds more weakly to astrophysics through backreaction.
- Star formation rate density: SFRD responds strongly to Ωm, σ8, and ASN1 across most redshifts, while ASN2 matters mainly at z < 4.AAGN1 and AAGN2 have little effect in IllustrisTNG but clearly affect SIMBA at z < 2.
- Star formation rate density: SFRD trends with ASN1 reverse between IllustrisTNG and SIMBA at lower redshifts, consistent with different feedback implementations and possible late-time wind recycling.At z > 3, increasing ASN1 decreases SFRD in both suites; at z ≲ 1.5, SIMBA shows the opposite trend.
- Machine-learning applications: CAMELS combines simulation outputs with machine learning to provide cosmology- and astrophysics-dependent predictions for key observables.The project contains thousands of simulations and models, enabling applications such as nonlinear interpolation and parameter estimation.
5. Output: SFRD from z=0 to z=7; 100 numbers
A neural network predicts the SFRD trajectory from six cosmological and astrophysical parameters. It captures the general trend accurately, but cannot reproduce realization-specific high-frequency variability caused mainly by cosmic variance.
- Performance: 0.12 dex is the SIMBA prediction error, compared with 0.106 dex for IllustrisTNG.The higher SIMBA error is attributed to noisier, less smooth SFRD curves.
- Performance: The network captures the general SFRD trend well across four held-out test simulations.The plotted predictions are compared with simulations excluded from training and validation.
- Limitations: Cosmic variance produces high-frequency SFRD variability that the network does not capture because the input lacks the initial density field.The approximately 20% cosmic-variance scatter sets a lower bound on the achievable error for this setup.
- Scope: The emulator is presented as a potential tool for fast parameter-space exploration, but the demonstration omits observational uncertainties and other ingredients needed for rigorous analysis.Its data are sampled equally in redshift, whereas many physical quantities evolve in physical time.
5.2. Constraining parameters
The inverse application uses SFRD measurements to infer cosmological and astrophysical parameters with a neural network. It constrains four parameters but cannot determine the two AGN-feedback parameters because their effects on SFRD are weak.
- Method: The network maps 100 SFRD values from z = 0 to z = 7 to six cosmological and astrophysical parameters.The data use the SIMBA LH set and are divided into 700 training, 150 validation, and 150 test simulations.
- Results: Ωm, σ8, ASN1, and ASN2 are constrained, whereas AAGN1 and AAGN2 cannot be determined from SFRD.The weak SFRD response to the AGN parameters explains their lack of constraint in this application.
- Results: The respective parameter errors are 0.055 for Ωm, 0.051 for σ8, 0.55 for ASN1, and 0.25 for ASN2.These are average errors over test-set realizations.
- Limitations: The constraints are not particularly tight because the simulations use a small (25 h^-1 Mpc)^3 volume and different random seeds, producing significant cosmic-variance scatter.The reported parameter errors therefore reflect substantial realization-to-realization variation.
- Limitations: The reported parameter error is an average, and estimating errors at each parameter-space point would ideally require multiple SFRD realizations with identical cosmology and astrophysics.Bayesian neural networks are suggested as an alternative when such repeated simulations are unavailable.
5.3. Symbolic regression
Symbolic regression uses genetic programming to derive compact analytic expressions for SFRD as a function of redshift and selected cosmological and astrophysical parameters. These expressions are less accurate than neural networks but offer interpretable parameter dependences.
- 5.3. Symbolic regression: Genetic programming searches combinations of functions, variables, and constants, retaining better-performing expressions across generations.The implementation uses Eureqa with operators including +, −, ×, ÷, log, exp, and ab.
- 5.3. Symbolic regression: The regression input is redshift, Ωm, σ8, ASN1, and ASN2, while the output is log10(SFRD).
- 5.3. Symbolic regression: Three compact expressions achieve errors of 0.19, 0.18, and 0.16 on both training and test sets.The expressions were selected for low training error and compactness; longer expressions are barely more accurate, while shorter ones have significantly lower error.
- 5.3. Symbolic regression: The neural network achieves a δ = 0.106 error on the same simulations, outperforming the analytic expressions.Expression accuracy depends on redshift; Eq. 20 is the most accurate and least redshift-dependent, whereas Eq. 18 is the least accurate overall.
- 5.3. Symbolic regression: Analytic expressions expose parameter interactions, such as ASN1 acting as a normalization constant in Eq. 19 and adding redshift dependence in Eqs. 18 and 20.The authors note that analytic forms can help explain dependencies even when they perform worse than neural networks.
5.4. Generative models
The generative-model application trains GANs on projected IllustrisTNG temperature fields to generate new maps with similar visual and statistical properties. Generated images transition smoothly in latent space and match real-map statistics well, without strong evidence of mode collapse.
- 5.4. Generative models: The training data comprise more than 135,000 augmented 64 × 64 temperature-map images extracted from projected simulation slices.Maps cover voids, filaments, galaxy groups, and feedback-induced bubbles across a broad astrophysical parameter range.
- 5.4. Generative models: GANs learn to generate samples from a low-dimensional manifold containing cosmic-web structures and varied cosmological and astrophysical parameters.The generator creates images while the discriminator distinguishes real from fake images.
- 5.4. Generative models: Generated temperature maps are visually difficult to distinguish from real maps and reproduce detailed features associated with star-forming gas around massive halos.
- 5.4. Generative models: Latent-space interpolations produce smooth, realistic transitions across many tested image pairs, providing no evidence of severe mode collapse.Mode collapse would make interpolated images vary non-smoothly.
- 5.4. Generative models: The generated maps match real maps in temperature power-spectrum and PDF statistics, including agreement of roughly 25% across almost four orders of magnitude in temperature.
5.5. Dimensionality reduction
Autoencoders reduce projected temperature fields to a lower-dimensional representation and test whether that representation generalizes across cosmologies and astrophysical models. Reconstructions remain accurate across varied simulations, though bottleneck size controls the trade-off between fidelity and compression.
- 5.5. Dimensionality reduction: Autoencoders are used as nonlinear dimensionality-reduction methods that learn a bottleneck representation from which temperature maps can be reconstructed.
- 5.5. Dimensionality reduction: The autoencoder is trained on IllustrisTNG CV temperature fields at fixed cosmology and astrophysics, then evaluated on maps from varied models.
- 5.5. Dimensionality reduction: The maximum reconstruction error is around 1.3 × 10−3, while the error distribution peaks around 5 × 10−4.
- 5.5. Dimensionality reduction: Maps from different cosmologies and astrophysics models are reconstructed with similar accuracy to the training-domain maps, although their error distributions differ slightly in peaks and tails.
- 5.5. Dimensionality reduction: A 500-neuron bottleneck is about 10% of the original image size, while reducing it to 100 neurons makes reconstructions much blurrier.Larger bottlenecks improve reconstruction but provide less compression.
- 5.5. Dimensionality reduction: Autoencoders can identify anomalies because unusual inputs produce deviations in reconstruction error.
6.1. Scientific goals
CAMELS aims to connect cosmological and astrophysical modeling with machine learning by providing theory predictions, marginalizing baryonic effects, and quantifying galaxy-formation dependencies. Its goals also include mapping N-body to hydrodynamic simulations and calibrating subgrid parameters against observations.
- 6.1. Scientific goals: CAMELS provides theory predictions for statistics or fields as functions of cosmology and astrophysics.
- 6.1. Scientific goals: The project seeks to extract cosmological information while marginalizing over baryonic effects.
- 6.1. Scientific goals: CAMELS aims to find the mapping between N-body and hydrodynamic simulations.
- 6.1. Scientific goals: The project quantifies how galaxy formation and evolution depend on astrophysics and cosmology.
- 6.1. Scientific goals: Machine learning is used to efficiently calibrate subgrid parameters in hydrodynamic simulations against observations.
6.2. Simulations
CAMELS comprises 4,233 simulations spanning cosmological and astrophysical parameter variations, multiple hydrodynamic and N-body codes, and controlled simulation subsets.
- Simulation suite: 4,233 simulations use N-body, AREPO, and GIZMO calculations in periodic (25 h^-1Mpc)^3 boxes, with IllustrisTNG and SIMBA subgrid physics.Hydrodynamic simulations evolve 256^3 dark matter particles and 256^3 fluid elements.
- Parameter space: Six parameters vary across CAMELS: Ω_m, σ_8, and four parameters controlling stellar and AGN feedback.
- Simulation subsets: The LH set contains 1,000 simulations with varied cosmological, astrophysical, and initial-seed values.
- Simulation subsets: The 1P, CV, and EX sets isolate one-at-a-time parameter changes, cosmic variance from initial seeds, and extreme feedback models, respectively.They contain 61, 27, and 4 simulations.
- Simulation suite: The hydrodynamic suites total 2,184 simulations, while their dark-matter-only counterparts provide 2,049 N-body simulations.
6.3. Cosmological and astrophysical properties
CAMELS spans twelve cosmological and astrophysical observables, revealing broad agreement between IllustrisTNG and SIMBA for many galaxy properties but important differences in baryonic effects.
- Observable coverage: Twelve quantities include matter and gas power spectra, halo and stellar mass functions, star formation, halo baryon fractions, temperatures, and galaxy scaling relations.
- Cross-model comparison: IllustrisTNG and SIMBA produce roughly similar mass functions, star formation rate density, and galaxy/halo scaling relations across the sampled parameter space.
- Cross-model comparison: Cosmic variance contributes significantly to overlap in halo temperature, galaxy scaling relations, and large-scale matter and gas power-spectrum amplitudes.
- Cross-model comparison: IllustrisTNG and SIMBA make significantly different predictions for halo baryon fractions and baryonic effects on the matter power spectrum.
- Baryonic interplay: Feedback-parameter changes can produce different or opposite observable responses under the IllustrisTNG and SIMBA model implementations.The global star formation rate history responds differently to ASN1 variations between the models.
6.4. Machine learning applications
CAMELS is designed for machine learning across a densely sampled six-dimensional parameter space, with applications spanning prediction, inference, generation, compression, and anomaly detection.
- Design and scope: More than 1,000 simulations per code/model sample the six-dimensional parameter space for machine learning.
- Prediction and inference: Neural networks predict cosmic star formation rate density with ≃30% accuracy, compared with ≃20% average scatter from cosmic variance.
- Prediction and inference: Likelihood-free neural-network inference estimates Ω_m, σ_8, ASN1, and ASN2 with average errors of 0.055, 0.051, 0.55, and 0.25, respectively.The reported errors arise from the small simulation volumes.
- Symbolic regression: Symbolic-regression equations predict star-formation rate density with ≃45% accuracy, versus ≃30% for neural networks and ≃20% intrinsic cosmic-variance scatter.
- Data generation: GAN-generated temperature maps agree with simulations within 15% for power spectra and 25% for PDFs.
- Compression and anomalies: Autoencoders reconstruct temperature fields across different cosmological and astrophysical parameters with training-set accuracy and identify non-background regions of the CAMELS logo as anomalies.
- Computational trade-off: Testing neural networks is computationally negligible, whereas autoencoder training required ∼150 GPU hours compared with ∼6,000 CPU hours for one hydrodynamic simulation.
- Computational trade-off: Required accuracy depends on the task: covariance generation may not need percent-level accuracy, while observable emulation may require sub-percent accuracy.
6.5. Limitations and extensions
CAMELS has resolution, volume, and parameter-scope limitations that constrain some cosmological applications, while planned extensions target higher resolution, larger volumes, and broader parameter coverage.
- Resolution: CAMELS cannot resolve scales below ∼1 h^-1kpc, and halos below 6.5 × 10^9(Ω_m−Ω_b)/0.251 h^-1M⊙ lack 100 dark matter particles.This limits probes relying on very small-scale matter structure, such as Milky Way sub-halos.
- Volume: The (25 h^-1Mpc)^3 volume omits long-wavelength modes important for large objects and proper matter-power-spectrum normalization across scales.Larger-volume and separate-universe simulations are planned extensions.
- Parameter scope: CAMELS varies only two cosmological and four astrophysical parameters, with Ω_b fixed when Ω_m varies, preventing separation of effects involving Ω_b/Ω_m.Future plans include varying h, n_s, M_ν, and w in larger-volume simulations.
- Overall scope: CAMELS is the largest such suite for machine learning but remains limited for cosmological-data analysis in different regimes.The current suite primarily supports proof-of-concept development for broader CAMELS goals.
6.6. Data access
CAMELS provides extensive simulation snapshots, catalogues, and derived data products for analysis. The project also makes analysis codes and technical details available online.
- CAMELS contains 143,922 full snapshots from 4,233 simulations.
- Each snapshot includes a halo/galaxy catalogue and additional products such as power spectra and bispectra.
- The CAMELS website provides data-access information, analysis codes, and further technical details about the simulations.
A. GAN ARCHITECTURE
The GAN uses convolutional layers to build generated images through progressively increasing spatial resolutions, while the discriminator architecture is introduced separately.
- Generator: The generator expands representations from 4 x 4 to 64 x 64 through successive ConvTranspose2d layers.The listed generator stages increase channels from 512 to 1 while raising spatial resolution.
- Generator: The generator includes ReLU activations between transposed-convolution stages.The architecture also lists LeakyReLU activations with parameter 0.2.
- Discriminator: The discriminator uses successive Conv2d layers and outputs the probability that an image is real.Its listed convolutional stages reduce spatial resolution toward a 1 x 1 output.
B. AUTOENCODER ARCHITECTURE
The autoencoder architecture combines convolutional encoding with transposed-convolutional decoding. Its listed layers progressively reduce and then restore spatial resolution.
- Encoder: The encoder uses Conv2d layers that reduce spatial dimensions from 32 x 32 to 1 x 1.The listed stages progress through 16 x 16, 8 x 8, and 4 x 4 representations before the bottleneck.
- Decoder: The decoder uses ConvTranspose2d layers to expand representations from 4 x 4 through 8 x 8, 16 x 16, and 32 x 32.
- Decoder: The decoder ends with a ConvTranspose2d layer producing a 1-channel 64 x 64 output.