Source-linked AI summary

The MultiDark Database: Release of the Bolshoi and MultiDark Cosmological Simulations

Kristin Riebe, Adrian M. Partl, Harry Enke, Jaime Forero-Romero, Stefan Gottloeber, Anatoly Klypin, Gerard Lemson, Francisco Prada, Joel R. Primack, Matthias Steinmetz, Victor Turchaninov

arXiv:1109.0003v2astro-ph.COastro-ph.IMcs.DB

TL;DR

Large cosmological simulations are difficult to distribute and analyze because their data volumes and formats overwhelm conventional access. This paper presents the MultiDark Database, which organizes two 8.6-billion-particle simulations in relational tables and exposes them through SQL queries. The first release enables catalogue, profile, and selected raw-particle access, while future releases are planned to add further snapshots, galaxy mocks, merger trees, and larger simulations.

  • Problem

    Large simulation datasets are difficult to distribute and analyze because their volume makes complete retrieval and local management impractical.

  • Method

    The paper organizes simulation data in a relational database and provides SQL-based server-side filtering, query access, and scientific analysis.

  • Results

    The first release provides access to two 8.6-billion-particle simulations, halo and subhalo data, profiles, and selected raw dark-matter particle data.

  • Takeaways & Limitations

    The database makes complex simulation results accessible through relational and Virtual Observatory-oriented infrastructure for community use and comparison.

  • Takeaways & Limitations

    The initial release has limited snapshot coverage and does not yet include the planned galaxy mock catalogues, additional merger trees, or larger-volume simulation data.

Abstract

from arXiv · show

We present the online MultiDark Database -- a Virtual Observatory-oriented, relational database for hosting various cosmological simulations. The data is accessible via an SQL (Structured Query Language) query interface, which also allows users to directly pose scientific questions, as shown in a number of examples in this paper. Further examples for the usage of the database are given in its extensive online documentation (www.multidark.org). The database is based on the same technology as the Millennium Database, a fact that will greatly facilitate the usage of both suites of cosmological simulations. The first release of the MultiDark Database hosts two 8.6 billion particle cosmological N-body simulations: the Bolshoi (250/h Mpc simulation box, 1/h kpc resolution) and MultiDark Run1 simulation (MDR1, or BigBolshoi, 1000/h Mpc simulation box, 7/h kpc resolution). The extraction methods for halos/subhalos from the raw simulation data, and how this data is structured in the database are explained in this paper. With the first data release, users get full access to halo/subhalo catalogs, various profiles of the halos at redshifts z=0-15, and raw dark matter data for one time-step of the Bolshoi and four time-steps of the MultiDark simulation. Later releases will also include galaxy mock catalogs and additional merging trees for both simulations as well as new large volume simulations with high resolution. This project is further proof of the viability to store and present complex data using relational database technology. We encourage other simulators to publish their results in a similar manner.

1. Introduction

Modern cosmological simulations produce enormous dark-matter datasets that are difficult to distribute and analyze in heterogeneous formats. The MultiDark Database addresses this challenge by combining relational storage with server-side SQL filtering, building on the Millennium Database approach.

  • Motivation: Dark-matter-only simulations require theoretical methods such as semi-analytical models or Halo-Occupation-Distribution approaches to predict galaxy properties.
  • Motivation: Billions of particles make simulation data difficult to distribute, especially when even compact catalogs become impractical for broad release.Different data formats further burden researchers who analyze simulation results.
  • Database approach: Server-side SQL filtering lets users analyze subsets and retrieve only results instead of managing multi-terabyte simulation datasets locally.The paper identifies server-side filtering as necessary for disseminating such large datasets.
  • Database approach: The MultiDark Database adopts the Millennium Database’s technology, implementation, and data structures to facilitate joint use and comparisons of halo statistics.The approach also aligns the data with Virtual Observatory-oriented access.
  • Paper scope: The paper presents database design, simulation data, storage details, example science cases, and halo-finder and merger-tree methods.

2. The MultiDark Database and its design

The MultiDark Database organizes simulation objects and derived products in linked relational tables accessed through SQL. Its web interface supports interactive querying, scripted retrieval, and registered-user storage of results, while a public mini-database provides a test environment.

  • Relational design: Relational tables store simulation objects and derived products, with rows representing objects and columns representing their properties.The FOF catalogue, for example, has one row per FOF group.
  • Relational design: Foreign keys link rows across tables, allowing relationships such as FOF groups and their particles to be represented explicitly.
  • Querying: SQL filters data through table relationships and lets the database engine handle execution plans, I/O, looping, and optimization.This gives users a direct route from a scientific question to an executable query.
  • Querying: The database uses Microsoft’s T-SQL dialect and supports extensions such as stored procedures for frequently used query components.
  • Access: A public miniMDR1 database exposes a roughly (100 h^-1 Mpc)^3 subvolume for exploring capabilities and developing queries before registration.
  • Access: Interactive results can be viewed, plotted, or downloaded in formats such as CSV and VOTable, while registered users can store results privately for further use.The interface is designed for interactive use but browser rendering limits make stored results useful for larger outputs.
  • Access: Scripted access supports large query-result retrieval and statistical analysis through tools including wget, Topcat, and IDL.

3. Simulation Data

The first MultiDark release contains two complementary 8.6-billion-particle simulations and multiple halo catalogues organized for further analysis. It provides BDM and FOF products with distinct definitions, linked particle data, and an explicit limitation on FOF post-processing.

  • Simulation data: Two separate databases contain the Bolshoi and MultiDark Run1 simulations, which use roughly 8–10 billion particles in different volumes, cosmologies, and codes.
  • Simulation data: Bolshoi uses a (250 h^-1 Mpc)^3 volume with approximately 8.6 · 10^9 particles and ART code, while MDR1 uses a (1 Gpc h^-1)^3 cube with the same particle count.
  • Halo catalogues: The database provides BDM and FOF halo catalogues containing halo properties such as positions, velocities, masses, and radii.Catalogue descriptions and additional storage details are provided in later sections and appendixes.
  • Halo catalogues: BDM catalogues use two halo-radius definitions: virial overdensity for BDMV and ∆200 = 200 ρcrit for BDMW.The BDMW halos are smaller than corresponding BDMV halos because their defining density is higher.
  • Halo catalogues: Each BDM halo and subhalo is characterized by 23 parameters, including coordinates, velocities, and two mass measures.For distinct halos, the two masses typically differ by at most 1–2 percent; the difference is larger for subhalos.
  • FOF catalogues: FOF groups have unique particle assignments, enabling links between FOF catalogues and particle tables and supporting unique progenitor-descendant relations.
  • FOF catalogues: FOF groups received no binding or unbinding post-processing, although their particle positions and velocities remain available for user-defined post-processing.

4. Data in the MultiDark Database

The MultiDark Database organizes halo catalogues, profiles, merger trees, substructure trees, and simulation particles for efficient retrieval and analysis. Its structures support spatial and mass-based queries, halo-history studies, and access to raw particle data, with documented scope boundaries for some catalogues and operations.

  • Halo catalogues: Halo catalogues provide records for BDM and FOF objects, with calculated properties organized by linking length and catalogue type.FOF catalogues include several linking lengths, while BDM catalogues are stored as BDMV and BDMW tables.
  • Halo catalogues: Spatial grid indexes and Peano-Hilbert keys enable fast retrieval of halos or FOF groups from a specified spatial region.
  • Halo catalogues: Mass and snapshot columns can be retrieved or sorted efficiently, supporting calculation of halo mass functions and their evolution.
  • Halo profiles: BDM profiles use logarithmically spaced shells out to 2Rvir and provide density, circular velocity, and related properties for halos with more than 100 particles.Each profile row is linked to its halo through bdmId, allowing all radial entries for a selected halo to be retrieved.
  • Merger trees: Merger trees trace progenitors backward from z = 0 roots, with the most massive progenitor defining the main branch and depth-first identifiers encoding tree membership.Identifiers such as treeRootId, fofTreeId, descendantId, and mainLeafId support tree and main-branch queries.
  • Trees and simulation particles: Merger trees cover z = 0 halos with more than 200 particles and terminate when the main progenitor falls below 20 particles, while FOF substructure trees organize progressively smaller linking lengths.The MultiDark tree sample is constructed to a maximum redshift of z = 5.4, and the database also exposes complete raw particle snapshots with positions and velocities.

5. Examples of using the database

The database supports diverse cosmological analyses through SQL-accessible halo profiles, particle data, aggregation, and downloadable subsets. Examples demonstrate velocity-profile studies, particle-level post-processing, halo-finder tests, and halo mass-function calculations.

  • The database supports analyses of halo properties, abundance, clustering, assembly, environments, voids, cosmic structure, merger trees, and mock catalogs.
  • Example 1: Velocity function: Halo profiles enable measurements of radial velocities across halo masses, revealing increasing infall amplitudes for group- and cluster-sized halos.Galaxy-size halos show no infall, while massive halos exhibit significant infall.
  • Example 2: Access to particles: Complete snapshot particle data can be retrieved for selected objects, including a 6.9 × 10^15 h−1 M⊙ supercluster containing 791743 particles.
  • Example 2: Access to particles: Retrieved particle positions and velocities enable individual post-processing, density visualization, and application of independent halo finders.The example used a 130h−1 kpc density grid and compared projected density with AHF-identified objects.
  • Example 2: Access to particles: Only about 2% of all particles need downloading to analyze spherical halos with mvir > 10^15h−1 M⊙ via superclusters above the same mass threshold.Retrieving particles for 1000 such halos required 6 hours and 10 minutes and had to be split into individual queries.
  • Example 3: Halo mass function: SQL aggregation produces halo mass functions from FOF and BDM catalogs at redshifts z = 0, z = 1, and z = 3.

6. Summary

The paper presents the MultiDark Database as a relational facility for hosting and analyzing two large cosmological simulations through SQL and web access. Its first release enables scientific queries and interoperability, while future releases are planned to expand snapshots, mock catalogs, and simulation volume.

  • The first release provides web-accessible relational data from the 8.6-billion-particle Bolshoi and MultiDark Run1 simulations.
  • SQL queries support scientific analysis, while tables and relations facilitate comparisons and compatibility with International Virtual Observatory Alliance standards.
  • Future releases are planned to add raw data for more snapshots, galaxy mock catalogs, and at least one larger-volume simulation.

Appendix A. Bound Density Maximum (BDM) halofinder

The BDM halo finder identifies halos and subhalos from density maxima using spherical overdensity, binding, and profile-based procedures. It produces extensive structural and dynamical properties, while its inertia-tensor axis ratios require concentration-dependent correction.

  • Halo and subhalo identification: BDM identifies halos and subhalos from particle-density maxima using a top-hat filter, spherical overdensity masses, and iterative bound-particle selection.The typical filter contains Nfilter = 20 particles; subhalo iterations remove unbound particles until convergence or the particle threshold is crossed.
  • Halo and subhalo identification: Distinct-halo centers are assigned to overlapping spheres with the deepest gravitational potential, while peripheral regions may still overlap.Halo radius and mass depend on whether distinct halos overlap.
  • Implementation: BDM accelerates particle searches with two-level link-lists and uses partial ranking to locate most-bound or nearest particles.MPI domain decomposition and OpenMP parallelization support the computation.
  • Catalog properties: The catalog provides 23 parameters per halo or subhalo, including coordinates, velocities, total mass, bound mass, velocity measures, concentration, size, spin, and shape.The catalog distinguishes the mass of all particles inside the virial radius from the mass of gravitationally bound particles.
  • Catalog properties: Vrms describes particle motions and contributes to kinetic-energy and virial analyses, while Vcirc is obtained by maximizing sqrt(GM(<R)/R) over logarithmic radial bins.The Vcirc search begins with the first bin containing at least 5 particles and uses Δlog R = 0.01.
  • Shape measurements: The halo shape is derived by diagonalizing a modified inertia tensor, but the reported axis ratios are not corrected for the spherical measurement region.The paper states that correction depends on halo concentration and gives corrections for flattened NFW profiles.

Appendix B. Friends-of-Friends halofinder (FOF)

The FOF halo finder groups particles through a linking-length criterion and implements a parallel hierarchical algorithm based on minimum spanning trees. The resulting groups support bulk, mass, velocity, spin, and shape measurements, while remaining aspherical by construction.

  • FOF definition: FOF groups are defined by a single relative linking length based on the mean inter-particle distance.For a given linking length, the method uniquely defines particle clusters whose members are separated by less than the threshold.
  • Parallel implementation: The parallel FOF implementation uses minimum spanning trees and subgraph decomposition to process 2048^3-particle simulations with low memory requirements.The geometric features of FOF objects allow the computation to be distributed across subgraphs.
  • Group structure: FOF groups cannot intersect, so each particle maps uniquely to one group; smaller-linking-length substructures lie completely within their hosts.This property supports database links between FOF groups and their particles.
  • Catalog properties: The database records each FOF group's center of mass, velocity, particle count, total mass, velocity dispersion, spin, and inertia-tensor axis ratios.FOF groups are intrinsically aspherical, so circular velocity is not defined in the same way as for spherical halos.

Appendix C. Merger trees for FOF catalogues

FOF merger trees are built by matching groups across consecutive snapshots through shared particles, producing branches with one descendant per halo and potentially many progenitors. The correspondence is incomplete because halos can disappear or be temporarily bridged.

  • Tree construction: Merger trees are constructed for z = 0 FOF groups with more than 200 particles by comparing consecutive snapshots and identifying shared-particle progenitors.Candidate progenitors must share at least 13 particles with the descendant group.
  • Limitations: FOF-to-tree correspondence is not always one-to-one, and some halos are excluded when they fall below the 20-particle detection threshold or are affected by temporary particle bridges.A disappearing halo branch is cut, while bridge-induced splitting can stitch a less massive clump's history to the most massive branch.

Appendix D. Database tables - Overview

The database overview organizes MultiDark and Bolshoi data into named tables and relational connections. A relational diagram documents the available tables, keys, and relationship cardinalities.

  • Table inventory: Table D.4 lists the names and descriptions of tables in the MDR1, miniMDR1, and Bolshoi databases.The paper directs readers to the MultiDark website for the complete table overview.
  • Relational structure: The relational diagram shows available tables, connecting keys, relation purposes, and one-to-one, one-to-many, or many-to-many cardinalities.Only relevant table-object keys are shown in the simplified diagram.
Loading 1109.0003v2…