Source-linked AI summary
Rapid identification of pathogenic bacteria using Raman spectroscopy and deep learning
Chi-Sing Ho, Neal Jean, Catherine A. Hogan, Lena Blackmon, Stefanie S. Jeffrey, Mark Holodniy, Niaz Banaei, Amr A. E. Saleh, Stefano Ermon, Jennifer Dionne
TL;DR
Rapid bacterial identification is limited by day-scale laboratory testing and the high signal-to-noise ratios required for accurate Raman classification. This paper applies deep learning to noisy Raman spectra, achieving comparable or improved identification accuracy across more isolate classes despite one-second measurements and roughly tenfold lower SNRs.
Problem
Bacterial identification can take days, while accurate Raman classification typically requires high signal-to-noise ratios.
Method
The study applies state-of-the-art deep learning techniques to noisy Raman spectra for clinical bacterial identification.
Results
One-second measurements with SNRs an order of magnitude lower than typical reported spectra achieved comparable or improved accuracy across more isolate classes.
Takeaways & Limitations
The approach supports bacterial identification from noisy, low-SNR Raman spectra across more isolate classes than typical Raman identification studies.
Takeaways & Limitations
Larger datasets are needed to achieve treatment recommendations as fine-grained as those from culture-based methods.
Abstract
from arXiv · showhide
Rapid identification of bacteria is essential to prevent the spread of infectious disease, help combat antimicrobial resistance, and improve patient outcomes. Raman optical spectroscopy promises to combine bacterial detection, identification, and antibiotic susceptibility testing in a single step. However, achieving clinically relevant speeds and accuracies remains challenging due to the weak Raman signal from bacterial cells and the large number of bacterial species and phenotypes. By amassing the largest known dataset of bacterial Raman spectra, we are able to apply state-of-the-art deep learning approaches to identify 30 of the most common bacterial pathogens from noisy Raman spectra, achieving antibiotic treatment identification accuracies of 99.0$\pm$0.1%. This novel approach distinguishes between methicillin-resistant and -susceptible isolates of Staphylococcus aureus (MRSA and MSSA) as well as a pair of isogenic MRSA and MSSA that are genetically identical apart from deletion of the mecA resistance gene, indicating the potential for culture-free detection of antibiotic resistance. Results from initial clinical validation are promising: using just 10 bacterial spectra from each of 25 isolates, we achieve 99.0$\pm$1.9% species identification accuracy. Our combined Raman-deep learning system represents an important proof-of-concept for rapid, culture-free identification of bacterial isolates and antibiotic resistance and could be readily extended for diagnostics on blood, urine, and sputum.
Introduction · Results · Deep learning for bacterial classification from Raman spectra
The paper addresses slow culture-based bacterial diagnosis by combining Raman spectroscopy with a CNN designed for noisy, low-SNR one-dimensional spectra. Using large reference and clinical datasets, the system classifies bacterial isolates and shows promising accuracy for species and antibiotic-resistance identification.
- Introduction: Culture-based diagnosis can take days, delaying targeted antibiotics and contributing to unnecessary broad-spectrum treatment.The paper motivates rapid, culture-free methods to support earlier targeted antibiotic prescription and help mitigate antimicrobial resistance.
- Introduction: Raman spectroscopy can identify bacterial species and antibiotic resistance, but weak scattering and subtle spectral differences require high-SNR, long measurements.Raman scattering has an ∼10−8 scattering probability, so background noise can mask phenotype-specific spectral differences.
- Introduction: The study trains a convolutional neural network to classify noisy bacterial spectra by isolate, empiric treatment, and antibiotic resistance.The approach targets the large diversity of clinically relevant bacterial species, strains, and resistance patterns that require comprehensive datasets.
- Deep learning for bacterial classification from Raman spectra: The CNN uses 25 one-dimensional convolutional layers and residual connections, replacing pooling with strided convolutions to preserve spectral-peak locations.The authors empirically find that preserving exact peak locations improves model performance.
Identification of empiric treatments and antibiotic resistance
The method groups bacterial isolates by recommended empiric antibiotic treatment and achieves 97.0±0.3% average accuracy, outperforming logistic regression and SVM. A binary CNN also differentiates MRSA from MSSA with 89.1±0.1% accuracy and an ROC AUC of 0.953.
- Empiric treatment identification: 97.0±0.3% average accuracy was achieved by grouping isolates according to recommended empiric antibiotic treatment.The 30 isolates were arranged into treatment-based groupings and summarized in a confusion matrix.
- Empiric treatment identification: 93.3% and 92.2% accuracies were achieved by logistic regression and SVM, respectively, compared with 97.0±0.3% for the method.These results compare the empiric-treatment classification approaches.
- Antibiotic resistance identification: 89.1±0.1% identification accuracy was achieved by a binary CNN differentiating methicillin-resistant from methicillin-susceptible S. aureus isolates.This model was trained as a step toward culture-free antibiotic susceptibility testing using Raman spectroscopy.
- Antibiotic resistance identification: 0.953 ROC AUC characterized MRSA-versus-MSSA classification, indicating that positive examples were ranked more likely as MRSA than negative examples with probability 0.953.The ROC curve supports tuning the binary decision toward higher sensitivity when false negatives are more consequential.
Extension to clinical patient isolates
Fine-tuning on small clinical datasets substantially improved bacterial identification in patient isolates, reaching high accuracy with only 10 spectra per isolate. Performance also improved on a second clinical dataset, while MRSA/MSSA classification remained comparatively limited.
- Clinical isolate identification: Using leave-one-patient-out cross-validation, the model was fine-tuned on 10 randomly sampled spectra from each clinical patient isolate and evaluated with 10 held-out spectra.The clinical dataset comprised 25 patient isolates, with five isolates from each of five empiric treatment groups.
- Clinical isolate identification: 99.0±1.9% species identification accuracy was achieved after clinical fine-tuning, compared with 89.0±3.6% for the reference-pretrained CNN baseline.The reported ± values are standard deviations across 10,000 sampling trials.
- Clinical isolate identification: 99.0% correct identification was obtained using 10 cellular spectra, within 1% of the 100.0% achieved with 400 spectra.The result supports accurate identification from the low bacterial-cell numbers expected in uncultured patient samples.
- Antibiotic susceptibility identification: 65.4±6.3% accuracy was achieved by the fine-tuned MRSA/MSSA classifier, compared with 61.7±7.3% for the pretrained binary classifier.This proof-of-concept used five additional clinical MRSA isolates and the same leave-one-patient-out process.
- Robustness across clinical datasets: 99.7±1.1% treatment group identification accuracy was achieved using only 10 spectra per patient on a second clinical dataset.Performance improved for both S. aureus and P. aeruginosa, demonstrating potential for continuous model improvement.
Discussion
Deep learning applied to noisy Raman spectra enables rapid bacterial identification and treatment classification, with fine-tuning supporting extension to clinical isolates. The platform could generalize to other spectroscopic tasks and, with automation, enable culture-free diagnostic testing, although broader datasets are needed for finer-grained treatment recommendations.
- Clinical extension: A CNN trained on noisy Raman spectra can identify clinically relevant bacteria and empiric treatment, while fine-tuning extends it to new clinical settings using few clinical isolates.The authors propose continuous fine-tuning to evaluate and improve deployed models.
- Generalizability: The model requires minimal modification for other identification problems and spectroscopic techniques, including materials identification, nuclear magnetic resonance, infrared, and mass spectrometry.The discussion presents this as a possible extension of the current Raman-CNN approach.
- Performance: 1 s measurements produced SNRs an order of magnitude lower than typical reported bacterial spectra while achieving comparable or improved identification accuracy across more isolate classes.This demonstrates performance under substantially shorter acquisition times than typical spectra.
- Biofluid translation: A broad SERS dataset could allow CNN processing of blood, sputum, or urine samples in a few hours, despite SERS variability and limited reproducibility on cell samples.SERS can increase signal strength by several orders of magnitude, but its variability complicates reliable diagnostics.
- Generalizability: Raman spectroscopy may identify phenotypes without specially designed labels, supporting generalizability to new targets compared with other culture-free approaches.The discussion contrasts Raman’s labeling requirements with methods involving single-cell sequencing, fluorescence, or magnetic tagging.
- Limitations and future work: Finer-grained treatment recommendations require larger datasets spanning resistant and susceptible isolates, antibiotic susceptibility profiles, cell states, growth media, and growth conditions.Automated sample preparation and acquisition would be needed to collect such datasets and support clinical translation.
- Clinical translation: An automated Raman-CNN system could scan every cell in patient samples, recommend antibiotic treatment without culture, and enable targeted treatment within hours.The authors associate this potential with reduced healthcare costs, antibiotic misuse, antimicrobial resistance, and improved patient outcomes.
Methods · Dataset · Dataset variance
The study used reference, fine-tuning, test, and clinical Raman datasets spanning bacterial, yeast, and Candida isolates, including an isogenic MRSA/MSSA pair. High intra-sample variance motivated using many spectra per isolate to represent distributions and improve prediction.
- Dataset: The reference dataset included 30 bacterial and yeast isolates spanning Gram-negative, Gram-positive, and Candida species.The dataset included multiple isolates from each bacterial grouping.
- Dataset: The dataset included an isogenic Staphylococcus aureus pair differing by the methicillin-resistance mecA gene, representing MRSA and MSSA.The MRSA variant contained mecA, whereas the MSSA variant did not.
- Dataset: The reference training dataset contained 2000 spectra per reference isolate plus isogenic MSSA at three measurement times.
- Dataset: The reference fine-tuning and test datasets each contained 100 spectra per reference isolate.
- Methods: Measurement times for the reference fine-tuning, test, and second clinical datasets increased from 1 s to 2 s to maintain consistent SNR as optical efficiency degraded.
- Methods: Antibiotic susceptibility was assessed by mecA PCR genotyping followed by phenotypic testing on Microscan Walkaway and VITEK 2 instruments.
- Dataset variance: For 19 out of 30 isolates, spectra from at least one other isolate were more similar on average than spectra from the same isolate.For E. faecalis 2, eight other isolates showed this pattern.
Sample preparation · Raman measurements
Bacterial isolates were prepared under consistent conditions on gold-coated silica substrates, then measured by mapped Raman spectroscopy and processed to retain primarily monolayer spectra. Monolayer and single-cell measurements had comparable signal-to-noise ratios, while separate test preparation supported classification independent of batch effects.
- Sample preparation: Isolates were cultured daily on blood agar, stored sealed at 4°C for 20 minutes to 12 hours, and prepared consistently between samples.Storage-time variation did not produce spectral changes greater than strain or isogenic differences.
- Sample preparation: Test samples were prepared separately from training samples, supporting classifications that were not caused by batch effects.Clinical isolates were also prepared as separate samples under consistent conditions.
- Sample preparation: Samples were made by suspending 0.6 mg biomass in 10 µL sterile water, or 0.4 mg in 5 µL for Gram-positive species, then drying 3 µL on gold-coated silica.Gold was deposited as a 200 nm coating on pre-cleaned microscope slides, and samples dried for 1 hour before measurement.
- Raman measurements: Raman spectra were mapped across dried monolayer regions using 633 nm illumination at 13.17 mW, a 300 l/mm grating, and 1.2 cm−1 dispersion.A 100X 0.9 NA objective produced an approximately 1 µm spot, and silicon was used for wavenumber calibration.
- Raman measurements: Each map comprised 45x45 spots spaced 3 µm apart, with spectra individually background-corrected using a fifth-order polynomial fit.The spacing avoided overlap between spectra, and the correction used the subbackmod Matlab function.
- Raman measurements: Spectra likely arising from aggregates or multilayer regions were excluded by discarding the 25 highest-intensity spectra, including those above two standard deviations from the mean.Most spectra came from true monolayers and approximately one cell per diffraction-limited laser spot.
- Raman measurements: Monolayer measurements had SNRs of 2.5±0.7, comparable to single-cell measurements at 2.4±0.6, while enabling semi-automated generation of a large training dataset.Both monolayer and single-cell measurements were collected.
- Raman measurements: The analyzed spectral range was 381.98–1792.4 cm−1, and each spectrum was normalized to a minimum intensity of 0 and maximum intensity of 1.SNR was calculated from the total intensity range divided by the intensity range in a signal-free 20-pixel window.
CNN architecture & training details
The model uses a 26-layer ResNet-based CNN selected over MLP and other CNN architectures, with shared classifier designs for species and MRSA/MSSA tasks. Training combines pre-training, fine-tuning, validation-based model selection, and independent testing, with clinical isolates evaluated by leave-one-patient-out cross-validation.
- CNN architecture: The 26-layer ResNet-based CNN has an initial convolution, 6 residual layers, and a final fully connected classification layer.Each residual layer contains 4 convolutional layers, and shortcut connections support gradient propagation and stable training.
- Architecture selection: The ResNet-based architecture performed best among the tested ResNet-based, MLP, and CNN architectures.Architecture hyperparameters were selected by grid search using one training and validation split on the isolate classification task.
- Classifier design: The 30-isolate classifier outputs probabilities across 30 classes, while binary MRSA/MSSA classifiers retain the architecture but change the final layer’s class count.The maximum output probability is used as the predicted class.
- Training and evaluation: Training uses pre-training followed by fine-tuning, with validation accuracy for model selection and testing on independently cultured and prepared samples.Accuracies are reported across 5 randomly selected train and validation splits, and the binary MRSA/MSSA classifier follows the same procedure.
- Clinical validation: Clinical isolates are fine-tuned with leave-one-patient-out cross-validation across 25 patient isolates from 5 species, assigning patients to training, validation, and test sets by fold.Each fold uses 15 patients for fine-tuning, 5 for validation-based model selection, and 5 for testing.
Clinical identification data analysis
Patient-isolate identification used random subsampling of Raman spectra followed by majority-vote classification. Clinical-dataset error values were estimated across 10,000 random trials, with accuracy capped at 100%.
- 400 spectra are measured from each patient isolate, and 10 are randomly selected for classification.
- The most common class among the 10 spectral classifications determines each patient-isolate identification, with ties broken randomly.
- 10,000 random selections of 10 spectra define the clinical-dataset standard deviations, with an upper accuracy bound of 100%.
- The second clinical dataset uses 10 of 100 spectra per patient isolate and a model pre-trained on the reference dataset and fine-tuned on the first clinical dataset.
Baselines · Two-sample test of sample means
PCA-reduced logistic regression and support vector machine baselines were evaluated on bacterial classification tasks, while Welch’s two-sample t-tests assessed whether CNN accuracy differences were statistically significant. The baselines achieved substantially lower reported accuracies than the CNN, and both comparisons rejected equal means at the 1e-6 p-level.
- Baselines: PCA reduced the input dimension from 1000 to 20 for logistic regression and support vector machine baseline experiments.The value was selected near test-accuracy saturation on one training and validation split for the 30-isolate task.
- Baselines: Using the first 20 principal components decreased computation costs and increased accuracy by reducing data noise.
- Baselines: Grid search selected each model’s regularization hyperparameter by validation accuracy within every cross-validation fold, with corresponding test accuracy reported.
- Baselines: 57.5% and 56.8% were achieved by logistic regression and support vector machine, respectively, on the 30-class task using both reference datasets.
- Baselines: 89.0% and 88.3% were achieved by logistic regression and support vector machine, respectively, on the empiric treatment task using both reference datasets.
- Baselines: 75.7% and 74.9% were achieved by logistic regression and support vector machine, respectively, on the 30-class task using only the fine-tuning reference dataset.On the empiric treatment task, the corresponding accuracies were 93.3% and 92.2%, respectively.
- Two-sample test of sample means: Welch’s two-sample t-test tested whether differences in mean clinical accuracy between the CNN and each baseline were statistically significant.This test accommodates samples that may have unequal variances.
- Two-sample test of sample means: 102.9 and 88.3 were the t-statistics for CNN-versus-logistic-regression and CNN-versus-support-vector-machine comparisons, respectively, rejecting equal means at the 1e-6 p-level.The calculations used nCNN = nLR = nSVM = 10000, with means µCNN = 89.0, µLR = 81.8, and µSVM = 82.9.
Biological materials availability
The authors make unique bacterial isolates available upon reasonable request. The study also documents reference and clinical isolate collections, including 25 patient-derived isolates spanning five bacterial species.
- Unique isolates are available from the authors upon reasonable request.
- Reference isolates and empiric treatment groupings are documented in Supplementary Table 1.Treatment choices were informed by infectious-disease recommendations and susceptibility trends at Stanford Hospital and the Veterans Affairs Palo Alto Health Care System.
- 25 clinical isolates comprised five each of E. coli, E. faecalis, E. faecium, P. aeruginosa, and S. aureus.The experiment collected 400 spectra for each isolate from patient samples.