Source-linked AI summary
Inductive Biases in Field-Level Cosmological Inference from Galaxy Catalogs
James O. Baldwin, Shy Genel, Francisco Villaescusa-Navarro
TL;DR
The paper asks how machine-learning models can extract cosmological information from galaxy catalogs when observables and architectural inductive biases vary. Using CAMELS simulations, it compares Deep Sets and KAN-based variants with relational GNNs on positions and peculiar velocities. Velocity-only Deep Sets recover Ω_m across suites, KANs do not improve on MLPs, and GNNs better exploit positional information; applying these findings to surveys requires realistic observational validation.
Problem
The paper examines which galaxy observables and architectural assumptions enable robust extraction of cosmological information from high-dimensional galaxy catalogs.
Method
The study performs likelihood-free field-level inference of Ω_m from CAMELS galaxy catalogs using Deep Sets with MLPs or KANs and spatially relational GNNs, testing positions and peculiar velocities across simulation suites.
Results
Velocity-only Deep Sets recover Ω_m across simulation suites, KANs provide no significant accuracy improvement, and GNNs improve over Deep Sets when positional information is modeled relationally.
Takeaways & Limitations
Peculiar velocities can be effectively exploited by set-based models, while spatial geometry is most effectively used by architectures that explicitly encode galaxy-galaxy relations.
Takeaways & Limitations
The results use exact simulated peculiar velocities and therefore require validation under realistic velocity-measurement noise, selection effects, and survey geometry.
Abstract
from arXiv · showhide
We perform field-level likelihood-free inference of the matter density parameter $Ω_m$ from simulated galaxy catalogs using machine learning models with differing inductive biases. Using hydrodynamic simulations from CAMELS, we examine how observable choice and architecture govern cosmological information extraction. We consider galaxy positions and line-of-sight peculiar velocities, separately and jointly, and compare permutation-invariant Deep Sets, implemented with either multilayer perceptrons (MLPs) or Kolmogorov-Arnold Networks (KANs), to graph neural networks (GNNs), which explicitly encode spatial relations. We test in-distribution and out-of-distribution (OOD) performance across simulations with different subgrid galaxy-formation prescriptions. Deep Sets infer $Ω_m$ from velocities alone with mean relative errors of approximately $18\%$ in-distribution and $\sim25\%$ OOD, with KANs and MLPs achieving comparable performance. In contrast, the same set-based approach does not yield useful $σ_8$ predictions in either in-distribution or cross-suite tests. Adding positions does not improve Deep Sets, while GNNs infer $Ω_m$ with mean relative errors of about $10\%$ in-distribution and $10$--$17\%$ OOD. These results indicate that peculiar velocities provide the dominant source of $Ω_m$ information for set-based models in this setting, while spatial information is most effectively used by architectures that explicitly encode galaxy-galaxy relations. Because the velocity inputs are exact simulated peculiar velocities, applications to survey data will require validation under realistic velocity-measurement noise, selection effects, and survey geometry.
1. INTRODUCTION
The paper studies how observable choice and architectural inductive bias determine which cosmological information machine-learning models can extract from galaxy catalogs. It focuses on Ω_m inference from positions and peculiar velocities using set-based and relational architectures, with cross-suite robustness as a central criterion.
- Methodological context: Field-level likelihood-free inference operates directly on simulated galaxy catalogs instead of hand-crafted summary statistics or an analytic likelihood.The paper uses “field-level” for direct catalog operation and “likelihood-free” for simulation-trained mappings without evaluating an analytic likelihood.
- Motivation and scope: The study varies galaxy features and model architectures to identify which combinations can access cosmological information from simulated catalogs.It compares permutation-invariant Deep Sets, including MLP and KAN implementations, with GNNs that encode spatial relations.
- Observable choice: Peculiar velocities are examined because they trace the large-scale gravitational potential, while positions carry inherently relational information.This motivates testing whether set-based models can exploit positional information without explicitly representing spatial structure.
- Evaluation scope: The analysis evaluates generalization across hydrodynamic simulation suites with different galaxy-formation prescriptions using cross-suite out-of-distribution performance.Robustness is framed as maintaining predictive performance under changed conditions, with cross-suite OOD behavior as a central criterion.
- Contribution: The paper’s contribution is to clarify where cosmological information resides in galaxy catalogs and how inductive biases govern access to it, rather than introduce a new inference method.The stated goal is to hold the inference task fixed while varying features and architectures, with implications for future survey pipelines subject to realistic observational effects.
2. SIMULATIONS
The simulations use CAMELS hydrodynamic galaxy catalogs spanning multiple solvers, subgrid prescriptions, parameter samplings, and feature representations. Catalog inputs include galaxy positions, line-of-sight peculiar velocities, and an optional global galaxy-count feature.
- Simulation setup: All CAMELS hydrodynamic runs use periodic (25 h^-1 Mpc)^3 boxes containing 256^3 dark-matter particles and initially 256^3 gas resolution elements.The simulation-set table caption summarizes these common box and particle specifications.
- Simulation suites: The dataset draws on four galaxy-formation models with different hydrodynamics solvers and subgrid physics implementations.The suites are Astrid, IllustrisTNG, Simba, and Swift-EAGLE, with IllustrisTNG also contributing the SB28 set.
- Cross-suite interpretation: The four suites share parameter labels but encode physically distinct astrophysical processes, so corresponding astrophysical parameters are not directly comparable across suites.This distinction matters for interpreting cross-suite evaluations.
- Parameter sampling: The work combines Latin Hypercube and Sobol-sequence designs, with LH sets containing 1,000 simulations per suite and SB28 containing 2,048 simulations in a 28-dimensional space.SB28 is unique to IllustrisTNG and varies Ω_b across the set.
- Input features: Galaxy inputs include positions and the z-component of peculiar velocity, with positions represented relative to the box center and optionally supplemented by log(N), the logarithm of galaxy count.Position features are the scalar vector (d_i, α_i, β_i), and additional properties such as v_z are concatenated pointwise.
3. ARCHITECTURES
The architecture comparison contrasts permutation-invariant Deep Sets and DeepKANs with spatially relational GNNs. Deep Sets aggregate pointwise galaxy representations, whereas GNNs propagate information through proximity-defined galaxy relations.
- 3.1. Deep Sets: Deep Sets map each galaxy to a latent representation, sum across the unordered catalog, and transform the aggregate into a global Ω_m prediction.This supports variable-length inputs and permutation invariance while allowing uncertainty prediction.
- 3.3. Kolmogorov–Arnold Networks (KANs) and DeepKANs: KANs replace fixed MLP node activations with learned univariate transformations on edges, assembled through composition and summation.Their spline parameterization uses a grid whose resolution controls the representation of the learned transformations.
- 3.3. Kolmogorov–Arnold Networks (KANs) and DeepKANs: DeepKAN embeds KANs in both pointwise and post-aggregation Deep Set functions while preserving exchangeability through summation.Spline-based edge functions provide flexible univariate transformations, and multiplicative nodes can represent product structures.
- Architectural comparison: Greater univariate functional expressivity does not improve peculiar-velocity inference over MLP-based Deep Sets, indicating that architecture and input information matter more than the approximator form here.Accordingly, the study does not test a KAN-based GNN.
- 3.2. Graph neural networks: GNNs represent galaxies as nodes and use proximity-defined edges to propagate information from neighboring galaxies before global pooling.This relational inductive bias is designed to capture clustering, tidal environments, and large-scale structure correlations.
4. EXPERIMENTAL DESIGN
The experiments use mass-cut marginalization, suite-specific test procedures, and a two-moment loss to train models that predict parameters with uncertainties.
- The training data use an 80/10/10 training-validation-test split from z = 0 galaxy catalogs in the LH set.
- 10 mass cuts per simulation produce 10,000 total simulations after marginalizing over galaxies above a prescribed mass cut.The simulations are split into training, validation, and test sets.
- The same mass-cut procedure is applied to Simba, Swift-EAGLE, and SB28 IllustrisTNG test sets.Astrid is used as the primary training suite for reasons described in the experimental design.
- Models jointly predict parameter means and variances using a modified two-moment loss designed to learn heteroscedastic uncertainties.Replacing arithmetic sums with logarithmic sums stabilizes training and balances contributions across parameters.
- Optuna selects configuration-specific hyperparameters for Deep Sets with MLPs and KANs, including layer, width, residual, and architecture settings.
5. RESULTS
Deep Sets extract useful Ωm information mainly from peculiar velocities, but adding positions does not help unless an architecture explicitly models galaxy relations. GNNs improve inference by exploiting spatial structure, while velocity-only models do not reliably recover σ8.
- 5.1. Deep Sets: MLPs vs KANs using Peculiar Velocities: Deep Set velocity models typically achieve RMSE values of ∼0.055–0.081 and mean relative errors of ∼17%–26%, outperforming simple mean and log N baselines.The comparison indicates that these models use more than the training prior or global catalog abundance.
- 5.1. Deep Sets: MLPs vs KANs using Peculiar Velocities: MLP- and KAN-based Deep Sets perform comparably for velocity-only Ωm inference, including similar cross-suite degradation.The results support information content, rather than function-approximator choice, as the primary limitation in this regime.
- 5.1. Deep Sets: MLPs vs KANs using Peculiar Velocities: Velocity-only DeepKAN uncertainty estimates are near nominal on held-out Astrid catalogs but become overconfident on cross-suite TNG(SB) data.Aggregate coverage changes from f1σ = 0.656 and f2σ = 0.923 in-distribution to f1σ = 0.491 and f2σ = 0.835 under distribution shift.
- 5.2. Deep Sets with Spatial Information: Adding positions to Deep Sets does not consistently improve Ωm inference, with OOD relative errors remaining ∼20–24%.Positions-only Deep Sets show no recoverable cosmological signal and perform comparably to models using only galaxy counts.
- 5.3. GNNs and Spatial Information: GNNs improve Ωm inference across in-distribution and cross-suite tests by encoding local spatial relations between galaxies.Relative to permutation-invariant Deep Sets, GNNs improve cross-suite R2, reduce relative errors, and produce regression slopes closer to one-to-one.
- 5.3. GNNs and Spatial Information: Positional information becomes useful when represented through graph relations, indicating that cosmological signal is encoded in spatial patterns rather than isolated coordinates.GNNs can model dependencies between nearby galaxies, whereas Deep Sets aggregate independent pointwise embeddings.
- 5.4. σ8 Inference: The velocity-only DeepKAN fails to recover a useful σ8 signal on both Astrid and cross-suite tests.Its predictions remain weakly correlated with the true values, so successful Ωm inference does not imply recovery of all velocity-field cosmology.
6. DISCUSSION
Peculiar velocities retain information about Ωm through their gravitational dynamics, whereas the tested learned summaries do not recover σ8 in this finite-volume setting. Spatial information improves inference only when the architecture explicitly represents galaxy relationships, and the idealized velocity input limits direct survey applicability.
- Peculiar velocities and cosmological parameters: Ωm sensitivity arises because it controls the gravitational sourcing and normalization of peculiar velocities through the potential and matter-growth dynamics.The discussion connects the Poisson, Euler, and continuity equations to coherent velocity generation across linear and nonlinear regimes.
- Peculiar velocities and cosmological parameters: The tested velocity summaries retain useful Ωm information but fail to recover σ8, with the latter result plausibly dominated by the limited CAMELS volume and missing long-wavelength modes.The finite (25 h−1 Mpc)^3 boxes provide limited sampling of the 8 h−1 Mpc scale relevant to σ8 and omit longer-wavelength contributions to coherent flows.
- Spatial information and architecture: Adding positions does not substantially improve Deep Sets, indicating that positional information is not effectively used when spatial coordinates are treated as independent features under global aggregation.Clustering information depends on separations and local environments, which require explicit relational modeling rather than pointwise aggregation.
- Spatial information and architecture: GNNs improve inference when positions are included, because message passing can capture spatial correlations that permutation-invariant Deep Sets do not represent explicitly.The improvement is attributed to relational inductive bias and the ability to model pairwise and higher-order structure through spatially defined edges.
- Scope and limitations: The results are not a direct forecast for surveys because the analysis uses exact simulated line-of-sight velocities instead of noisy, indirectly inferred measurements with selection and coverage effects.Observed peculiar-velocity catalogs inherit distance uncertainties, calibration systematics, selection effects, and incomplete sky coverage.
7. CONCLUSION
The study finds that peculiar velocities provide substantial Ωm information for set-based models, while relational architectures better exploit spatial information. KANs do not improve accuracy over MLPs, and survey applications remain contingent on realistic validation.
- Velocity-only models recover Ωm across simulation suites, but the same setting does not recover a robust σ8 signal.The velocity-based information is therefore more effective for constraining Ωm than σ8 in this analysis.
- KAN components do not significantly improve inference accuracy over alternative pointwise approximators, shifting attention toward input information and architectural bias.The result suggests that univariate function-approximation flexibility is not the limiting factor in this regime.
- GNNs improve performance relative to Deep Sets when positional information is incorporated through relational modeling.This is attributed to spatial relationships and anisotropic correlations that permutation-invariant architectures do not explicitly represent.
- Architectural inductive bias governs how cosmological information is extracted: velocities work effectively with set-based models, whereas spatial information benefits from relational GNNs.The findings emphasize matching model structure to the physical organization of the input field.
- Applications to real surveys require validation with noisy velocity measurements, calibration systematics, selection effects, and survey geometry.The analysis uses exact simulated line-of-sight peculiar velocities rather than realistic observational measurements.
A. TRAINED ON TNG
Models trained on the LH set of IllustrisTNG with line-of-sight peculiar velocities were evaluated across several simulation suites. Their overall behavior remained broadly consistent with models trained on Astrid despite somewhat degraded in-distribution performance.
- TNG-trained velocity-only Deep Set models show broadly consistent cross-suite behavior with Astrid-trained models, despite somewhat degraded in-distribution performance.The models use either MLP or DeepKAN architectures and are evaluated on Astrid, Simba, Swift-EAGLE, and the SB set of IllustrisTNG.
B. COMPARISON WITH PREVIOUS WORK
The comparison examines alternative implementations and summarizes their evaluation settings. The MetaLayer-based model struggles to extract the Ωm signal, with the discrepancy localized to the layer architecture rather than the input features or training suite.
- The comparison uses models trained on peculiar velocities and reports RMSE, R2, PCC, relative error, and reduced χ2 across multiple simulation suites.Table 4 covers TNG-trained Deep Set models, while Table 5 covers ASTRID-trained models using the CosmoGraphNet framework.
- The implementation replaces fully connected Linear layers with KAN layers in the DeepKAN variant.
- The MetaLayer architecture struggles to extract the underlying Ωm signal compared with the Linear-based implementation.The results are presented in Figure 8 and Table 5.