Source-linked AI summary
Accelerated Nuclear Magnetic Resonance Spectroscopy with Deep Learning
Xiaobo Qu, Yihui Huang, Hengfa Lu, Tianyu Qiu, Di Guo, Tatiana Agback, Vladislav Orekhov, Zhong Chen
TL;DR
NMR spectroscopy can require long acquisition times, motivating reconstruction from limited measurements. The paper introduces DL NMR, trained solely on synthetic FID/spectrum pairs, and reports high-quality reconstruction across protein spectra, while omitting a feedback connection when fully sampled FIDs are unavailable in practice.
Problem
NMR spectroscopy is useful in chemistry and biology but often requires long experimental times, creating a need for reliable reconstruction from limited data.
Method
DL NMR trains a neural network solely on synthetic FID/spectrum pairs and uses dense CNN processing with data consistency for reconstruction.
Results
DL NMR achieves reconstructed spectra quality comparable to low-rank reconstruction, with high-fidelity small-protein reconstructions having R2 > 0.99 and reliable results from 30% and 10% NUS data.
Takeaways & Limitations
The results support fast, robust NMR spectrum reconstruction from limited measurements without requiring realistic spectra for network training.
Takeaways & Limitations
The feedback connection is discarded during reconstruction because fully sampled FIDs are unavailable in practice.
Abstract
from arXiv · showhide
Nuclear magnetic resonance (NMR) spectroscopy serves as an indispensable tool in chemistry and biology but often suffers from long experimental time. We present a proof-of-concept of application of deep learning and neural network for high-quality, reliable, and very fast NMR spectra reconstruction from limited experimental data. We show that the neural network training can be achieved using solely synthetic NMR signal, which lifts the prohibiting demand for a large volume of realistic training data usually required in the deep learning approach.
manuscript.
The paper presents DL NMR as a neural-network approach for reconstructing spectra from undersampled FIDs and documents its methodology, data availability, and implementation context.
- Resources: Synthetic FID training data and several experimental spectra are made available through cited repositories or the corresponding author.The paper also states that code is available upon reasonable request.
- Method: The method is implemented in separate training and prediction phases using paired FID and spectrum data.The training phase learns from computer-simulated undersampled FIDs and target spectra.
- Method: DL NMR maps undersampled FID signals to target spectra through a trained neural network.The processing uses a neural-network mapping from input FIDs to spectra.
1.1 Training phase
The training pipeline generates synthetic fully sampled and undersampled FID/spectrum pairs, processes an artifact-laden initial spectrum with dense CNN and data consistency, and optimizes network parameters across reconstruction stages.
- 1.1.1 Generate the fully sampled spectrum and the undersampled FID: Synthetic training uses fully sampled FIDs generated from exponential signal models and Poisson-gap undersampling operators.The simulation varies amplitudes, phases, decay times, and frequencies to generate training examples.
- 1.1.1 Generate the fully sampled spectrum and the undersampled FID: 40000 FID/spectrum pairs are simulated for neural-network training.Each pair combines an undersampled FID with its corresponding fully sampled spectrum.
- 1.1.2 Generate the initial spectrum from the undersampled FID: The initial spectrum is formed by zero-filling unacquired FID positions, producing strong artifacts before neural-network processing.The adjoint undersampling operator and forward Fourier transform are used in this initialization.
- 1.1.3 Reduce spectrum artifacts and 1.1.4 Enforce data consistency: A dense CNN reduces spectral artifacts, while data consistency aligns reconstructed values with acquired FID data.The dense CNN contains eight convolutional layers, and the consistency module balances predicted and acquired sampled points.
- 1.1.4 Enforce data consistency and 1.1.5 Loss function and trained optimal parameters: Repeated CNN and data-consistency stages progressively improve spectral quality before the network output is optimized over training pairs.The loss optimization trains network parameters across reconstruction stages.
1.2 Reconstruction phase
During reconstruction, DL NMR maps an undersampled experimental FID directly to a reconstructed spectrum using the trained network function.
- 1.2 Reconstruction phase: The trained network maps an undersampled FID to a reconstructed spectrum during prediction.The reconstruction is represented as a function of the undersampled FID and trained parameters.
- 1.2 Reconstruction phase: The feedback connection is discarded during reconstruction because a fully sampled FID is unavailable in practice.This defines a practical boundary between the training and reconstruction configurations.
- 1.2 Reconstruction phase: The reconstruction phase is presented alongside comparisons with low-rank and compressed-sensing approaches.For 2D NMR, compressed sensing is excluded because low rank had previously been shown to outperform it.
2.1 Experiments Setup
The experiments evaluate DL NMR on multiple 2D and 3D protein spectra, using non-uniform sampling and varied acquisition conditions, dimensions, and sampling fractions.
- Experiment design: The study includes four 2D and four 3D spectra from small, large, and intrinsically disordered proteins.Direct dimensions were processed with NMRPipe before reconstruction.
- Sampling: Non-uniform sampling tables are generated with Poisson-gap sampling to reduce data-acquisition time.NUS denotes spectrometer acquisition in non-uniform sampling mode.
- 3D spectra: The 3D experiments include HNCO spectra of azurin, MALT1, and alpha-synuclein, plus an HNCACB spectrum of GB1-HttNTQ7.The datasets span different protein sizes, spectrometers, probes, and spectral dimensions.
- Sampling conditions: The 3D datasets use sampling fractions including 30% for MALT1 and 15% for alpha-synuclein.The corresponding fully sampled spectra have sizes 1024×57×70 and 1024×64×64, respectively.
2.2 Reconstructed 2D HSQC Spectrum of CD79b
DL NMR reconstructs the CD79b 2D HSQC spectrum from limited NUS data with quality comparable to LR at 25% sampling and improved robustness and correlation at lower sampling rates.
- The comparison uses LR as a representative NUS reconstruction method and evaluates reconstructions from limited experimental data.The section presents DL NMR as a deep-learning reconstruction approach for NMR spectra.
- At 25% NUS, DL NMR achieves the same reconstructed-spectrum quality as LR for CD79b.Both methods have peak-intensity correlation values approaching 0.9999, with peak shapes close to the fully sampled spectrum.
- At 10% and 15% NUS, DL NMR provides higher correlation values and lower dispersion across 100 NUS trials.These results indicate more stable reconstruction under reduced sampling.
2.3 Other 2D Spectra Reconstruction
DL and LR both reconstruct several 2D protein spectra accurately from 25% NUS data, while DL shows higher intensity correlations when fewer data are available.
- The experiments demonstrate applicability of the trained neural networks to three additional spectra.These include two HSQC spectra and one TROSY spectrum.
- Both DL and LR achieve peak-intensity correlations above 0.98 at 25% NUS across additional 2D spectra.The evaluated spectra include ubiquitin HSQC, GB1 HSQC, and ubiquitin best-TROSY spectra.
- At 25% NUS, both methods produce peak shapes almost identical to the fully sampled spectra.The comparisons include the ubiquitin HSQC, GB1 HSQC, and ubiquitin best-TROSY datasets.
- With fewer data, DL outperforms LR in intensity correlations, indicating higher acceleration factors for data acquisition.Figure S2-6 reports Pearson correlation coefficients over 100 NUS resampling trials.
2.4 3D Spectra Reconstruction
DL reconstructs 3D protein NMR spectra with high fidelity across fully sampled and experimentally undersampled datasets, including large and intrinsically disordered proteins. The method remains effective at 10% NUS in the reported examples.
- Evaluation scope: The 3D evaluation compares DL with state-of-the-art CS across small, large, and intrinsically disordered proteins.The datasets include azurin, GB1-HttNTQ7, MALT1, and alpha-synuclein.
- 3D reconstruction fidelity: Both DL and CS produce 3D reconstructions close to fully sampled spectra, with R2 > 0.99 for peak-intensity correlations.The high-fidelity comparisons use azurin and GB1-HttNTQ7 spectra reconstructed from undersampled data.
- Large protein: For MALT1, DL reconstructs HNCO spectra well from 30% NUS and remains similar to the 30% reconstruction at 10% NUS.The 10% dataset was created by randomly selecting one of every three points from the 30% NUS data.
- Computational evaluation: Reconstruction timing is evaluated after common direct- and indirect-dimension processing is omitted from the comparison.The computational setup includes a Tesla K40M GPU for DL and 24 CPU threads for LR.