Source-linked AI summary
Single-particle cryo-electron microscopy: Mathematical theory, computational challenges, and opportunities
Tamir Bendory, Alberto Bartesaghi, Amit Singer
TL;DR
This article examines the algorithmic and mathematical challenges of reconstructing 3-D molecular structures from cryo-EM data. It surveys computational approaches and discusses theoretical results, including conditions under which molecular structure is uniquely determined from image data.
Problem
Cryo-EM reconstruction involves computational challenges from experimental images, including increasing computational burden at higher resolution and uncertainty about suitable dimensionality-reduction techniques.
Method
The article introduces cryo-EM algorithms and mathematical problems, including dimensionality reduction, optimization, and statistical models for reconstruction-related tasks.
Results
Under stated assumptions, the molecular structure is uniquely determined from experimental images up to orthogonal matrices, and the third moment can determine the structure.
Takeaways & Limitations
The article connects cryo-EM reconstruction challenges with mathematical analysis of identifiability and computational methods.
Takeaways & Limitations
Common-lines estimation cannot work when the underlying object is a 2-D image rather than a 3-D object.
Abstract
from arXiv · showhide
In recent years, an abundance of new molecular structures have been elucidated using cryo-electron microscopy (cryo-EM), largely due to advances in hardware technology and data processing techniques. Owing to these new exciting developments, cryo-EM was selected by Nature Methods as Method of the Year 2015, and the Nobel Prize in Chemistry 2017 was awarded to three pioneers in the field. The main goal of this article is to introduce the challenging and exciting computational tasks involved in reconstructing 3-D molecular structures by cryo-EM. Determining molecular structures requires a wide range of computational tools in a variety of fields, including signal processing, estimation and detection theory, high-dimensional statistics, convex and non-convex optimization, spectral algorithms, dimensionality reduction, and machine learning. The tools from these fields must be adapted to work under exceptionally challenging conditions, including extreme noise levels, the presence of missing data, and massively large datasets as large as several Terabytes. In addition, we present two statistical models: multi-reference alignment and multi-target detection, that abstract away much of the intricacies of cryo-EM, while retaining some of its essential features. Based on these abstractions, we discuss some recent intriguing results in the mathematical theory of cryo-EM, and delineate relations with group theory, invariant theory, and information theory.
I. INTRODUCTION
Single-particle cryo-EM has rapidly advanced from limited, low-resolution applications to abundant near-atomic structures, while posing a difficult low-SNR inverse problem that demands substantial computational and mathematical tools.
- The cryo-EM reconstruction problem: Cryo-EM reconstructs a molecule’s 3-D electrostatic potential from 2-D particle projections recorded as micrographs.The target is a high-resolution estimate of the molecular structure, including atomic-scale features when resolution permits.
- The cryo-EM reconstruction problem: The central difficulty is estimating 3-D structure when particle orientations and locations are unknown in a very low-SNR regime.Low contrast without heavy-metal stains and high noise from radiation-limited electron doses compound the estimation challenge.
- Technological development: Since 2013, single-particle cryo-EM has undergone rapid transformations, producing abundant structures approaching atomic resolution.Direct electron detectors improved image quality and SNR, while high frame rates enabled motion correction from recorded movies.
- Motivation: Cryo-EM mitigates limitations of X-ray crystallography by avoiding crystallization, preserving near-physiological conformations, and recording particles independently.Independent particle projections can support estimation of multiple structures associated with different functional states.
- Article scope: The article introduces cryo-EM’s algorithmic challenges, leading software packages, and mathematical and statistical principles.It also presents multi-reference alignment and multi-target detection as abstract models for broader analysis, and suggests future research directions.
II. THE CRYO-EM RECONSTRUCTION PROBLEM
Single-particle cryo-EM reconstructs a 3-D molecular structure from noisy 2-D projections with unknown orientations and shifts. Its workflow moves from movie preprocessing and particle extraction through classification, model building, refinement, and post-processing.
- Imaging model: Each particle image is modeled as a rotated, projected, shifted, PSF-convolved version of the 3-D structure plus noise.The tomographic projection integrates along the z-axis, and the resulting image is sampled on a Cartesian grid.
- Imaging model: The inverse problem estimates the 3-D structure from 2-D images while treating particle rotations and shifts as nuisance variables.These variables are unknown but are not themselves the reconstruction target.
- Intrinsic ambiguities: Reconstruction is possible only up to global rotation, molecular location, and handedness ambiguities.Handedness cannot be determined from cryo-EM images alone because a structure and its reflection produce identical projection sets.
- Computational pipeline: The computational pipeline preprocesses movies, estimates CTFs, picks and optionally prunes particles, classifies images, builds initial models, refines volumes, and post-processes structures.Some workflows apply steps iteratively, such as repeating particle picking with class-derived templates.
- Structural variability: Structural variability changes the task from estimating one volume to estimating a distribution of conformations, often modeled discretely or on a low-dimensional space.A 3-D classification step can cluster projection images into different structural conformations.
- Noise model: Noise is commonly modeled as Gaussian after movie-frame averaging, although its spectrum is non-white and is often approximated as radially symmetric.At the frame level, shot noise follows a Poisson distribution; practical methods estimate noise-spectrum parameters during reconstruction or from signal-free regions.
III. MAIN COMPUTATIONAL CHALLENGES
Cryo-EM computation is difficult because measurements combine extreme noise, missing viewing information, and massive high-dimensional data. These conditions undermine direct alignment and reconstruction, requiring statistical and Fourier-based methods that avoid or manage unknown variables.
- Core challenges: Cryo-EM data combines high noise, missing viewing directions and locations, and massive datasets with high-dimensional reconstruction variables.Typical 200×200×200 volumes contain 8 million voxels, while movie datasets can reach several Terabytes.
- High noise: When σ →∞, the estimated rotation becomes uniformly distributed, and the Cramér–Rao bound remains proportional to σ2 and independent of N.Thus collecting more measurements does not make individual rotations accurately estimable at sufficiently high noise.
- High noise: The same high-noise effect makes alignment against many noiseless projection templates produce no salient correlation peak.As noise increases, the alignment distribution over SO(3) becomes flatter.
- Missing data: Known viewing directions reduce reconstruction to a linear inverse problem governed by the Fourier slice theorem, but cryo-EM samples Fourier space non-uniformly.Coverage of viewing directions affects solution quality, limiting direct use of classical filtered back projection.
- Missing data: Common-lines can estimate viewing directions, but any such estimator is expected to fail at low SNR.Because the reconstruction target is the 3-D structure rather than the viewing directions, methods such as maximum likelihood and the method of moments can circumvent rotation estimation.
- Massive datasets and dimensionality: High resolution compounds the burden because more parameters require more data while computation becomes severely more expensive.Modern datasets also require early attention to storage and computational throughput.
IV. 3-D RECONSTRUCTION FROM PROJECTIONS
3-D reconstruction methods alternate between assigning projection directions and updating the molecular structure. Hard assignment uses best matches, whereas soft assignment integrates uncertainty through maximum-likelihood or Bayesian models, especially ML-EM.
- Hard and soft refinement: Hard-assignment refinement alternates between assigning each image a best projection direction and reconstructing the structure by linear inversion.Fourier-space implementations exploit fast projection computation and diagonal CTF operators.
- Hard and soft refinement: Soft-assignment methods assign similarity scores or responsibilities to multiple image–projection pairs according to a generative statistical model.These methods are variants of expectation-maximization and support maximum-likelihood or Bayesian formulations with priors.
- ML-EM framework: ML-EM estimates the structure while marginalizing over nuisance transformations rather than jointly estimating every transformation for every observation.Joint estimation causes the parameter count to grow with the number of measurements and can make maximum likelihood inconsistent.
- ML-EM framework: In cryo-EM, the nuisance transformation rotates and projects the volume, then applies translation and PSF convolution over a 5-D space of rotations and translations.The transformation distribution is generally unknown and is estimated as part of the model.
- ML-EM framework: Each ML-EM iteration computes an E-step over nuisance variables and an M-step that updates the structure, often by solving a linear system under Gaussian assumptions.The posterior is nondecreasing across iterations, but non-convexity prevents a guarantee of reaching the MAP estimator.
- Practical refinement strategies: Frequency marching begins from a smooth low-resolution volume and gradually incorporates higher-frequency content during refinement.The associated prior initially regularizes high frequencies strongly, then increases their variance as the estimated structure gains detail.
- Practical refinement strategies: Common statistical priors are computationally convenient but may bias reconstruction, motivating biologically informed priors as an open challenge.The paper specifically points to previously reconstructed structures as one possible source of biological knowledge.
B. Ab initio modeling
Ab initio methods construct initial cryo-EM models without requiring an initial guess. They typically produce low-resolution estimates that can later be refined, while addressing the importance of avoiding model bias.
- B. Ab initio modeling: Ab initio models are constructed without an initial guess and usually provide low-resolution estimates for later refinement.The section motivates robust initialization methods by emphasizing model bias as a crucial issue in cryo-EM and statistics.
1) Model bias and validation:
Cryo-EM validation must distinguish data-supported structure from artifacts introduced by non-convex optimization, initialization, and model bias. Template-based detection can systematically steer reconstructions toward assumed structures.
- Validation: Non-convex cryo-EM reconstruction can produce outputs that depend on algorithm initialization.This dependence motivates testing convergence from many starting points.
- Validation: Validation asks whether a reconstructed 3-D model faithfully represents the underlying data, using biological plausibility, computational tests, repeated initializations, and cross-researcher comparison.Similar outcomes across multiple initializations strengthen confidence, while public data enable independent examination.
- Model bias: Model bias can arise when improper particle-picking templates cause erroneous detections that bias 3-D reconstruction toward the selected templates.The paper identifies template matching as a source of systematic error in particle picking.
- Model bias: The “Einstein from noise” experiment shows that aligning and averaging pure-noise images can produce an image resembling the reference rather than zeros.This demonstrates how alignment can bias results toward a chosen reference even without underlying signal.
- Model bias: Without prudent algorithmic design, reconstructed structures may reflect researchers’ presumptions rather than the structure that best explains the data.The paper presents this as the practical lesson of model-bias experiments.
- Optimization: Stochastic gradient descent reduces per-iteration cost by using randomly selected images or subsets to approximate the full gradient.Its randomness is also commonly believed to help escape local minima in some non-convex problems.
3) Estimating viewing directions using common-lines:
Common-lines methods infer relative viewing directions from Fourier-domain line agreements between projection images. Synchronization improves noise tolerance, while method-of-moments analysis links identifiable moments to sample complexity and computational cost.
- Common-lines: For generic asymmetric structures, Fourier transforms of projection pairs agree along identifiable central common-lines.Cross-correlating candidate central lines can locate the shared line.
- Common-lines: Three common-line pairs between three images uniquely determine their relative rotations, forming the basis of angular reconstitution.The method determines two Euler angles from a pair and the third using a third image.
- Synchronization: Synchronization estimates all image rotations jointly from pairwise common-lines, improving noise tolerance compared with sequential assignment.Spectral algorithms and semidefinite programming provide computational approaches with theoretical guarantees.
- Limitations: In the low-SNR regime, rotation estimation fails; synchronization therefore requires high-SNR images, often obtained through 2-D classification.Because classification blurs fine details, the resulting method has limited resolution.
- Limitations: A nontrivial object symmetry creates multiple common-lines between image pairs that must be considered.The common-lines method also cannot determine relative angles when the underlying object is only 2-D.
- Method of moments: The method of moments estimates parameter-dependent moment tensors by averaging observations, with one-pass estimation potentially suitable during acquisition.In cryo-EM, these moments depend on the volume, viewing-direction distribution, translations, and noise.
- Method of moments: The first moment distinguishing parameters determines low-SNR sample complexity, while lower-order moments reduce the number of equations and computational complexity.The identifying moment also determines the estimation rate.
- Method of moments: Under the stated assumptions, low-SNR sample complexity scales as σ4 for non-uniform rotations and σ6 for uniform rotations.The non-uniform case is identifiable from the second moment, whereas the uniform case requires the third moment.
5) Frequency marching:
Frequency marching progressively increases reconstruction resolution by estimating viewing directions and solving for higher-frequency structure from a low-resolution model. Motion correction addresses beam-induced drift before frames are combined.
- Frequency marching: Frequency marching expands a reconstruction from low to high resolution through progressively added frequency shells.The process may be explicit through an ab initio model or implicit through iteration-dependent regularization.
- Frequency marching: At each iteration, experimental images are compared with simulated projections of the current estimate and assigned viewing directions.The updated structure is then computed by solving a linear least-squares problem for another frequency shell.
- Motion correction: Beam-induced sample motion mitigates high-resolution information during data collection, motivating motion correction.Motion correction aligns and averages frames to partially reduce motion blur and improve signal-to-noise ratio.
- Motion correction: Because drift is not homogeneous across a micrograph, local motion-correction techniques were developed alongside global strategies.Per-particle correction tracks individual particles throughout electron exposure, while semi-local methods use grid-defined subregions.
- Motion correction: MotionCor2 estimates patchwise motion by cross-correlation, fits a time-varying 2-D polynomial, and sums corrected frames with optional dose weighting.The polynomial is quadratic in frame coordinates and cubic in time.
- Motion correction: Per-particle motion-correction strategies have been reported to yield significant resolution improvements.Figure 6 illustrates examples of different motion-correction strategies.
B. CTF estimation and correction
CTF correction must address microscope-induced frequency attenuation and missing or unstable frequencies. The workflow estimates CTF parameters from Thon rings and either phase-flips images, applies Wiener filtering, or incorporates CTF effects into reconstruction.
- CTF effects: The CTF alters image frequencies through the microscope PSF, with defocus trading low-resolution contrast against high-resolution detail.High defocus enhances low-resolution features and SNR, whereas low defocus enhances fine details at lower contrast.
- CTF estimation: CTF estimation fits modeled parameters to Thon rings in the micrograph power spectrum.CTFFIND4 and Gctf generate templates and select the one that best matches the measured rings.
- CTF correction: Direct CTF inversion is difficult because many small values and zero-crossings suppress frequencies.The microscope parameters determining these crossings are difficult to control and must therefore be estimated from acquired data.
- CTF correction: Phase flipping corrects Fourier-space phases by retaining only the sign of the CTF and ignoring amplitude changes.Wiener filtering attempts to correct both amplitudes and phases but requires SNR knowledge and cannot recover missing frequencies.
- CTF correction: Incorporating CTF effects into the reconstruction forward model can make full correction viable across micrographs with different PSFs.Frequencies suppressed in one micrograph may appear in another.
- Particle picking: Particle picking seeks to detect and extract noisy 2-D projections, but high-resolution reconstruction typically requires hundreds of thousands of particles.Manual picking is time-consuming and can introduce subjective bias.
- Particle picking: Automated particle picking can use reference windows, correlation, and a support-vector-machine classifier, while template-based methods risk systematic bias.Picked sets also contain non-informative particles such as noise, partial particles, or adjacent particles.
D. 2-D classification
2-D classification groups experimental images into in-plane-rotation-invariant class averages to improve signal quality, but low SNR, scale, and model mismatch make the task computationally difficult. Steerable PCA reduces the dimensionality-reduction cost by exploiting rotational symmetry.
- Classification goals: 2-D classification groups images from similar viewing directions into class averages after appropriate in-plane alignment.These averages support ab initio modeling, particle picking, particle assessment, and removal of non-informative particles.
- Computational challenges: Low SNR impedes accurate clustering and alignment, while hundreds of thousands of images make optimal clustering computationally challenging.Choosing an appropriate image-similarity metric is also unresolved.
- ML-EM: ML-EM models images using K class averages, in-plane rotations, translations, an estimated PSF, and Gaussian noise.The model assumes uniformly distributed classes and in-plane angles, isotropic Gaussian translations, and ε_i ∼ N(0, σ^2I).
- ML-EM limitations: ML-EM repeatedly processes all experimental images, assumes only K possible class averages despite continuous orientations, and tends to produce few low-resolution classes.The latter behavior reflects a winner-takes-all effect in which images favor higher-SNR classes.
- Alternative methods: A nearest-neighbor alternative averages a few properly aligned images using rotationally invariant bispectra, while steerable PCA compresses images before processing.The bispectrum remains unchanged under in-plane rotation, enabling rotation-invariant comparisons.
- Steerable PCA: O(NL^4 + L^6) becomes O(NL^3 + L^4) in steerable PCA because rotational integration makes the covariance block diagonal, reducing nonzero entries from L^4 to O(L^3).The first reduced term computes the sample covariance over N images; the second performs eigendecomposition over blocks.
VI. MATHEMATICAL FRAMEWORKS FOR CRYO-EM DATA ANALYSIS
The paper introduces multi-reference alignment and multi-target detection as abstractions for analyzing cryo-EM mathematically and statistically. These models connect reconstruction to sample complexity, invariant moments, and direct reconstruction from micrographs.
- Multi-reference alignment: Multi-reference alignment abstracts observations as transformed versions of an unknown signal with unknown group actions and Gaussian noise.Recovery estimates the signal's orbit, so the signal is identifiable only up to a group transformation.
- Multi-reference alignment: At high SNR, estimating group elements, undoing their actions, and averaging yields sample complexity proportional to 1/SNR.The averaging variance is proportional to 1/(N · SNR).
- Sample complexity: At low SNR, the MRA sample complexity is proportional to 1/SNR^q̄, where q̄ is the first moment distinguishing different signals.When N · SNR^q̄ remains bounded above in the asymptotic regime, every method has estimation error bounded away from zero.
- Sample complexity: For cryo-EM with perfect particle picking and no CTF, uniform rotations give sample complexity 1/SNR^3, whereas generic non-uniform rotations require 1/SNR^2.The uniform case is determined by the third moment, while the non-uniform case is determined by the second moment.
- Invariant theory: The role of moments motivates method-of-moments algorithms and invariant-polynomial methods, including the third-order bispectrum used in classification and ab initio modeling.An invariant polynomial is unchanged under the relevant group action.
- Multi-target detection: Multi-target detection models repeated unknown signals at unknown locations in noisy measurements, with signal-location detection corresponding to particle picking.At low SNR, detection and particle picking are precluded, motivating direct signal estimation.
- Multi-target detection: Under certain generative models, the signal can be estimated at any noise level from third-order bispectrum statistics without explicitly estimating signal locations.Numerical experiments suggest the bispectrum can also estimate multiple signals from blind-deconvolution mixtures.
- Cryo-EM modeling: An extended multi-target-detection model treats projections of a 3-D molecule as random signals embedded in micrographs and supports direct reconstruction without intermediate particle picking.The model includes rotations, tomographic projection, and the microscope point-spread function; a full analysis remains lacking.
VII. REMAINING COMPUTATIONAL AND THEORETICAL CHALLENGES
The section surveys unresolved computational challenges in cryo-EM, including noisy detection, structural heterogeneity, and reliable validation. It highlights opportunities for mathematical and computational advances across these problems.
- Multi-target detection: At high noise levels, reliable detection of repeated signal occurrences becomes challenging, motivating direct signal estimation without intermediate detection.Multi-target detection abstracts particle picking, while averaging is effective when noise is low.
- Heterogeneity: Structural heterogeneity is ill-posed because N 2-D images do not contain enough information to recover N 3-D structures without additional assumptions.The section distinguishes discrete models with a few volumes from continuous models based on low-dimensional structure.
- Heterogeneity: Discrete heterogeneity models extend maximum-likelihood expectation-maximization to multiple volumes but do not scale well for large K and ignore correlations between functional states.These drawbacks can cause the model to overlook information about relationships among states.
- Heterogeneity: Continuous heterogeneity models represent conformations in low-dimensional linear spaces, manifolds, or collections of moving rigid domains.Suggested tools include PCA and diffusion maps, but no consensus exists on the appropriate model or established computational tools.
- Validation: Cryo-EM lacks agreed computational verification methods, and current validation practices can be vulnerable to systematic errors from shared refinement procedures.Independent half-data reconstructions estimate resolution, while tilt-pair validation is useful only for intermediate-resolution structures.
- Validation: Developing confidence intervals for reconstructed structures that are immunized against systematic errors remains an important computational challenge.Existing validation also includes checking consistency across tilt-pair images and comparing reconstructions with structures from X-ray crystallography or NMR.
C. Theoretical foundations of cryo-EM
The section identifies open theoretical questions about cryo-EM sample complexity, computational efficiency, molecular-size limits, and machine-learning reliability. It presents abstract models and cross-disciplinary analysis as opportunities for progress.
- Open theoretical questions: Many mathematical and statistical properties of the cryo-EM problem remain unexplained, including its sample complexity and information-computational gaps.Having enough information to solve a problem does not imply that a polynomial-time algorithm exists.
- Open theoretical questions: For non-uniform rotation distributions, the second moment may suffice information-theoretically to estimate a 3-D structure, even when efficient computation remains unresolved.The section notes that higher-order moments may be needed to design computationally efficient algorithms in difficult regimes.
- Algorithmic theory: Theoretical understanding is also incomplete for algorithms such as ML-EM, the method of moments, and SGD, with information-computational gaps observed in MRA and MTD variants.These questions concern algorithm-specific properties beyond the existence of sufficient information.
- Molecular-size limits: The size limit of molecular structures accessible to cryo-EM remains an open question because small molecules produce low contrast and low micrograph SNR, hindering particle picking.A cited paper challenges this limitation by proposing direct reconstruction from micrographs without particle picking at any SNR level, in principle.
- Machine learning: Deep learning has been applied to particle picking, validation, reconstruction, pruning, heterogeneity analysis, and denoising, but its impact on 3-D reconstruction and heterogeneity remains limited.The section attributes supervised-learning concerns partly to model bias, because reconstructions can depend heavily on training data rather than experimental images.
- Cross-disciplinary opportunities: The article frames cryo-EM as an inverse problem linking signal processing, information theory, statistics, machine learning, and group theory.It also reviews multi-reference alignment and multi-target detection as abstract frameworks for studying essential cryo-EM features.