Source-linked AI summary
CryoBench: Diverse and challenging datasets for the heterogeneity problem in cryo-EM
Minkyu Jeon, Rishwanth Raghu, Miro Astore, Geoffrey Woollard, Ryan Feathers, Alkin Kaz, Sonya M. Hanson, Pilar Cossio, Ellen D. Zhong
TL;DR
Heterogeneous cryo-EM reconstruction lacks standardized ground-truth benchmarks and reliable metrics for comparing methods. CryoBench introduces five synthetic datasets, evaluation metrics, and analyses of state-of-the-art tools across diverse heterogeneity types and difficulty levels. Its overall benchmark quantification aligns well with qualitative observations, while future versions should address more realistic noise and combined heterogeneity.
Problem
Heterogeneous cryo-EM reconstruction lacks common benchmarks with ground truth and metrics suitable for reliable method evaluation and comparison.
Method
CryoBench combines five synthetic datasets, quantitative metrics, qualitative visualizations, and analyses of ten state-of-the-art heterogeneous reconstruction tools.
Results
Quantification of reconstruction performance aligns well with qualitative observations across CryoBench evaluations.
Takeaways & Limitations
CryoBench provides a shared resource for comparing heterogeneous reconstruction methods and motivating new development across cryo-EM and related imaging fields.
Takeaways & Limitations
Current benchmark extensions could add more realistic noise statistics, joint compositional and conformational heterogeneity, and non-structural sources of heterogeneity.
Abstract
from arXiv · showhide
Cryo-electron microscopy (cryo-EM) is a powerful technique for determining high-resolution 3D biomolecular structures from imaging data. Its unique ability to capture structural variability has spurred the development of heterogeneous reconstruction algorithms that can infer distributions of 3D structures from noisy, unlabeled imaging data. Despite the growing number of advanced methods, progress in the field is hindered by the lack of standardized benchmarks with ground truth information and reliable validation metrics. Here, we introduce CryoBench, a suite of datasets, metrics, and benchmarks for heterogeneous reconstruction in cryo-EM. CryoBench includes five datasets representing different sources of heterogeneity and degrees of difficulty. These include conformational heterogeneity generated from designed motions of antibody complexes or sampled from a molecular dynamics simulation, as well as compositional heterogeneity from mixtures of ribosome assembly states or 100 common complexes present in cells. We then analyze state-of-the-art heterogeneous reconstruction tools, including neural and non-neural methods, assess their sensitivity to noise, and propose new metrics for quantitative evaluation. We hope that CryoBench will be a foundational resource for accelerating algorithmic development and evaluation in the cryo-EM and machine learning communities. Project page: https://cryobench.cs.princeton.edu.
1 Introduction
Cryo-EM can capture biologically important structural heterogeneity, but heterogeneous reconstruction lacks standardized ground-truth benchmarks and suitable evaluation metrics. CryoBench addresses these gaps with challenging synthetic datasets, quantitative and qualitative metrics, and analyses of existing methods.
- Motivation: Cryo-EM images have low signal-to-noise ratios, unknown particle poses, and both conformational and compositional heterogeneity.These properties make 3D reconstruction difficult while preserving access to structural variability that structure prediction tools typically cannot capture.
- Research gap: Existing heterogeneous reconstruction research lacks common benchmarks and metrics suitable for comparing methods.The field has no benchmark analogous to MNIST or ImageNet for driving standardized progress.
- Research gap: Prior validation commonly used real data without ground truth, making reconstruction accuracy and the reliability of scientific conclusions difficult to assess.Earlier datasets included real structures, synthetic blob-like volumes, and pseudo-real motions, but these approaches provided limited standardized validation.
- CryoBench: CryoBench provides five challenging synthetic datasets spanning conformational and compositional heterogeneity, with ground-truth poses, states, and imaging parameters.The datasets range from interpretable diagnostic cases to harder settings intended to motivate new reconstruction methods.
- CryoBench: CryoBench evaluates state-of-the-art methods using metrics for both latent inference and volume reconstruction, supported by qualitative visualizations.The benchmark and evaluation tools are publicly available through the project website.
2 Background and Related Work
Heterogeneous cryo-EM reconstruction has expanded rapidly, but the field still lacks standardized, diverse benchmarks and reliable metrics. Existing datasets and evaluation practices often cover only particular heterogeneity types or lack ground truth.
- Heterogeneous reconstruction: Heterogeneous reconstruction methods differ in volume representation and inference strategy, including voxel grids, meshes, implicit neural representations, statistical inference, and gradient-based optimization.These approaches reflect broad methodological diversity in how structural variability is modeled and inferred.
- Past benchmarks: Common real datasets include ribosome assembly states for compositional heterogeneity and the pre-catalytic spliceosome for conformational heterogeneity.Such datasets have been reused for benchmarking despite the absence of a generally established benchmark suite.
- Past benchmarks: Simulated datasets provide ground truth but have generally been generated ad hoc for individual studies rather than as standardized benchmarks.Applying many methods to one dataset also fails to test performance across conformational motions and compositional changes.
- Metrics: Evaluation metrics for heterogeneous reconstructions are less straightforward than for homogeneous reconstruction, where FSC-based resolution assessment is mature.Standard FSC is flawed for heterogeneous reconstruction because it gives a global resolution assessment and is typically computed on independent half-sets as a self-consistency measure.
3 CryoBench Design
CryoBench constructs synthetic cryo-EM datasets by simulating images from ground-truth atomic models under a standard forward model. Its datasets span simple and challenging conformational motions, noise levels, molecular-dynamics trajectories, and compositional mixtures.
- Dataset construction: CryoBench generates ground-truth atomic ensembles and simulates the cryo-EM forward process to produce synthetic images.Density volumes are derived from atomic electron-scattering potentials before image formation.
- Image formation model: Each image is modeled as a CTF-modulated projection of a volume at a pose, with additive white Gaussian noise on a default 128 × 128 grid.The forward model explicitly controls imaging parameters needed for quantitative evaluation.
- Conformational heterogeneity: IgG-1D models a one-dimensional continuous circular motion by rotating an antibody dihedral angle, yielding 100 atomic models and 100,000 noisy projection images.The dataset uses a 3.6-degree sampling interval and an SNR of 0.01.
- Conformational heterogeneity: IgG-1D-noisier and IgG-1D-noisiest test robustness at SNRs of 0.005 and 0.001, respectively.These variants isolate the effect of increasing image noise on reconstruction methods.
- Conformational and compositional heterogeneity: IgG-RL introduces more complex conformational variability through random linker conformations and Fab-domain orientations, while Spike-MD uses molecular-dynamics models of SARS-CoV-2 spike motion at higher resolution.Ribosembly represents compositional heterogeneity across ribosome assembly states, and Tomotwin-100 scales compositional mixtures to 100 cellular complexes.
4 Evaluation Framework
CryoBench evaluates heterogeneous reconstruction with fixed-pose and ab initio methods using qualitative visualizations and quantitative metrics for embeddings, volumes, and poses. Its framework compares learned heterogeneity representations against ground truth and assesses reconstruction quality with per-image FSC-based measures.
- Evaluation setup: CryoBench evaluates seven fixed-pose methods and three ab initio variants across heterogeneous reconstruction tasks.The fixed-pose methods include 3D Class, 3DVA, 3DFlex, CryoDRGN, CryoDRGN-AI-fixed, Opus-DSD, and RECOVAR; ab initio methods are also evaluated.
- Datasets: Three datasets provide complementary tests: Spike-MD contains 46,789 structures, while Ribosembly contains 16 ground-truth assembly states.Spike-MD represents molecular-dynamics conformational motion; Ribosembly represents compositional heterogeneity across assembly states.
- Qualitative evaluation: Qualitative evaluation uses latent-embedding visualizations and representative density volumes to inspect learned heterogeneity and reconstructed structures.The framework visualizes method-specific latent distributions and samples representative volumes for visual inspection.
- Embedding Comparisons: Embedding comparisons use Neighborhood Similarity, Information Imbalance, and clustering metrics to quantify local-neighborhood preservation, shared information, and compositional-state separation.Neighborhood Similarity compares matching neighbors with ground truth; Information Imbalance compares feature spaces asymmetrically, and ARI/AMI compare cluster assignments with structural labels.
- Volume Metrics: Volume quality is summarized with Area Under the FSC Curve, while Per-Image FSC jointly evaluates conformation estimation and reconstruction quality.The AUC compares reconstructed volumes against ground truth and is reported for the Per-Image FSC curves in Table 1; pose error is additionally computed for ab initio methods.
5 Results
Across datasets, reconstruction difficulty increases with complex motions, many structures, and challenging pose inference. RECOVAR often performs strongly, while current methods struggle most on Spike-MD and Tomotwin-100.
- IgG-1D and IgG-RL: RECOVAR achieves the highest reconstruction quality on IgG-1D and superior FSC performance on the more challenging IgG-RL dataset.IgG-RL is harder because Fab orientation inference limits resolution of the moving Fab.
- IgG-1D: On IgG-1D, most methods preserve interpretable latent neighborhoods, but 3DFlex fails to capture the rotating Fab domain and ab initio methods fail at the lowest SNR.Volume metrics decrease as noise increases, as expected.
- Spike-MD: Spike-MD contains 46k unique structures, and methods show poor neighborhood overlap below 35% at small-to-intermediate radii.Different methods learn varied latent-space topologies, with neither ground-truth coordinate correlating well with the UMAP embeddings.
- Ribosembly: Ribosembly methods largely reconstruct the ground-truth states, but embeddings only achieve intermediate 50-80% neighborhood similarity and may mix assembly states.RECOVAR clusters best, while 3D Class and 3D Class abinit mix states within classes and have lower volume accuracy.
- Evaluation metrics: Table 2 evaluates clustering with ARI and AMI by k-means clustering particle embeddings using the number of ground-truth structures.The table defines the clustering evaluation used for comparisons such as Tomotwin-100.
- Tomotwin-100: Ab initio and 3D Classification-based methods fail to resolve Tomotwin-100, while current ab initio methods cannot recover its 100 ground-truth structures.Linear methods also have limited capacity for 200 kDa proteins, whereas cryoDRGN achieves almost perfect k-means clustering in the fixed-pose setting.
6 Conclusion and Future Directions
CryoBench combines five heterogeneous datasets, broad baseline analyses, and quantitative and qualitative evaluation tools. It offers both interpretable diagnostic tasks and challenging benchmarks, while identifying realism gaps for future extensions.
- Contributions: CryoBench provides five synthetic datasets, analyses of ten state-of-the-art tools, and metrics spanning representation learning, reconstruction quality, and end-to-end performance.Quantitative reconstruction results align well with qualitative observations.
- Future directions: The benchmark pairs simple datasets for interpretable development and validation with challenging datasets intended to push heterogeneous cryo-EM reconstruction.The authors anticipate these tasks may motivate work on distribution generalization, neural rendering, and biophysical priors.
- Limitations: The current benchmark uses simpler Gaussian noise and does not include more complex noise statistics or jointly simulated compositional and conformational heterogeneity.A multislice simulator was tested but omitted because it was much slower while producing qualitatively similar images and method performance.
- Future directions: Future benchmark versions could add custom metrics for molecular-dynamics data and non-structural heterogeneity such as junk particles and non-uniform pose distributions.The authors anticipate such developments will support applying novel methods to real-data biological discovery.
7 Data and Software Availability
CryoBench datasets, pose and CTF data, ground-truth resources, and evaluation scripts are made available through online repositories under stated licensing and deposition plans.
- Access: CryoBench datasets and tools are available from the project website.The project page is the primary access point for the released resources.
- Dataset deposition: Datasets are deposited to Zenodo, with downsampled images, CTFs, pose data, consensus volumes, and FSC masks included.Files are provided in .mrcs, .txt, .star, and Python pickle formats.
- Dataset deposition: Full-resolution images and ground-truth PDB files and volumes are planned for deposition to BioImage Archive.The datasets are provided under the Creative Commons Attribution 4.0 International license.
- Software: Scripts for simulating cryo-EM images and computing metrics are available in the CryoBench GitHub repository.
A Dataset Design
CryoBench constructs five synthetic datasets spanning designed, disordered-linker, molecular-dynamics, ribosome-assembly, and cellular-complex heterogeneity, with controlled imaging conditions and ground truth. The datasets vary in structural complexity, molecular composition, noise, and visual difficulty.
- Dataset scope: Five datasets cover conformational and compositional heterogeneity, ranging from designed antibody motions to ribosome assembly states and 100 cellular complexes.The collection includes IgG-1D, IgG-RL, Spike-MD, Ribosembly, and Tomotwin-100.
- IgG-1D: IgG-1D models a continuous circular motion by rotating a heavy-chain dihedral angle and sampling 100 structures at 3.6-degree intervals.Each conformation contributes 1,000 CTF-applied projection images, producing 100k images at SNR 0.01.
- IgG-RL: IgG-RL generates 100 random antibody-linker conformations by sampling backbone dihedral angles from disordered-peptide Ramachandran distributions while rejecting steric clashes.
- Complex heterogeneity: Spike-MD uses molecular-dynamics conformations of SARS-CoV-2 spike, while Ribosembly and Tomotwin-100 represent 16 ribosome states and 100 cellular complexes, respectively.Spike-MD contains 46,789 unique conformations sampled at an artificially high temperature; its images use D = 256 and SNR 0.1.
- Imaging conditions: Signal-to-noise ratio is defined as signal variance divided by noise variance, with Gaussian white noise added to reach the desired level.The entire D × D image is treated as signal, and example images are provided because SNR depends on the signal definition.
B Experimental Settings
The benchmark evaluates heterogeneous reconstruction methods under standardized masks, consensus volumes, latent representations, and dataset-specific settings. It combines neural, PCA-based, mesh-deformation, and discrete-classification approaches with FSC-based and clustering-based evaluation.
- Preprocessing: Masks are generated by aggregating ground-truth volumes and applying dataset-specific dilation and soft-padding procedures.Spike-MD instead uses the union of binarized volumes with cryoDRGN gen_mask; masks are shown in Figure S4b.
- Compared methods: The benchmark compares cryoDRGN, CryoDRGN-AI, Opus-DSD, RECOVAR, 3DFlex, 3DVA, and 3D Classification across the CryoBench datasets.These methods span implicit neural representations, convolutional networks, PCA, mesh deformation, probabilistic PCA, and discrete mixture models.
- Preprocessing: Consensus volumes for cryoSPARC methods are created by backprojecting all images with ground-truth poses, and the resulting volumes are shown in Figure S4a.
- Latent representations: Latent dimensionality varies by method and dataset, with class posterior lengths of 10 or 20 for most datasets and 20 for Spike-MD classification.Table S1 summarizes the number of latent dimensions used to model heterogeneity.
- Evaluation: Per-Conformation FSC is also evaluated with full and Fab masks, where overall performance decreases and 3DVA’s ranking changes relative to Opus-DSD.The result indicates that 3DVA’s performance largely stems from non-Fab regions.
- Evaluation: Sample FSC evaluates unsupervised representative volumes by matching each sampled volume against all ground-truth volumes and taking the maximum FSC-AUC.This avoids using ground-truth information for latent clustering, although correspondence to a specific ground-truth conformation is unavailable.
C.3 Additional Noise Comparison Results
Additional noise experiments show that increasing noise reduces reconstruction quality and makes conformational distinctions harder to recover, with detailed FSC curves supplied for comparison.
- Noise robustness: At SNR 0.005 and 0.001, volume metrics decrease and methods become less able to differentiate conformations compared with the original IgG-1D dataset.These conditions correspond to IgG-1D-noisier and IgG-1D-noisiest.
- FSC curves: Figure S10 provides all 100 masked FSC curves for IgG-1D, enabling per-conformation comparison across methods.
- FSC curves: Figure S11 summarizes average FSC curves across conformations with error bars showing standard deviation.
C.5 Additional Information Imbalance Results
Additional analyses examine how learned embeddings relate to ground-truth poses, CTF parameters, and molecular-dynamics collective variables. Most methods show limited imbalance for pose and CTF information, while Spike-MD neighborhoods remain poorly aligned at small radii.
- Pose and CTF: For pose and CTF parameters, most methods lie near the orthogonal information-imbalance region (1,1), while CryoDRGN and Opus-DSD show stronger pose entanglement.Opus-DSD and ab initio 3D Classification are generally farthest from the orthogonal region for CTF.
- Spike-MD: Spike-MD embeddings show relatively low neighborhood similarity to ground-truth molecular-dynamics collective variables at small neighborhood radii.This agrees with qualitative UMAP observations.
- Spike-MD: 3DVA lies on the shared information line at (0.5,0.5) for Spike-MD, paralleling its result on IgG-1D.
- Additional analyses: Figure S24 reports neighborhood similarity and information imbalance for Spike-MD, while Figure S6 compares masked and unmasked FSC distributions for IgG-1D.
C.7 Pose Error
CryoBench evaluates pose accuracy for ab initio heterogeneous reconstruction methods against ground-truth image poses. CryoDRGN-AI generally exhibits the lowest pose error across the evaluated datasets.
- Pose-error results: CryoDRGN-AI generally exhibits the lowest pose error among the evaluated ab initio methods.Pose-error distributions are shown for IgG-1D, IgG-RL, Spike-MD, Ribosembly, and Tomotwin-100.
- Evaluation scope: The analysis compares ab initio methods across datasets spanning antibody conformational motions, ribosome assembly, and diverse cellular complexes.For Tomotwin-100, each structure is aligned separately before pose errors are computed.
- Pose-error metric: Pose error is quantified by median rotation and translation errors after global reference-frame alignment.Rotation uses geodesic error in degrees, while translation uses L2 distance in pixels.
D.5 Image Acquisition and Analysis
Cryo-EM acquisition produces two-dimensional micrographs containing multiple particle images, which are boxed for analysis. Reconstruction then uses image series to generate three-dimensional voxelized volumes.
- Image acquisition: A micrograph is a two-dimensional electron-microscopy image containing a field of view with multiple particles.Micrographs may include temporal movie frames that are corrected for motion.
- Particle extraction: Individual particles are biomolecular structures or recorded measurements boxed from a larger micrograph.Particle images are typically processed as cropped regions from the wider field of view.
- Reconstruction: Reconstruction generates a three-dimensional voxelized volume from a series of two-dimensional images.The paper distinguishes homogeneous reconstruction of one volume from heterogeneous reconstruction of multiple structural states.