Source-linked AI summary
Machine learning-assisted directed protein evolution with combinatorial libraries
Zachary Wu, S. B. Jennifer Kan, Russell D. Lewis, Bruce J. Wittmann, Frances H. Arnold
TL;DR
Directed protein evolution faces costly experimental screening when multiple positions are mutated simultaneously. This paper uses machine learning to model sequence-function relationships and focus experimental testing on predicted combinatorial libraries. The approach improved fitness outcomes on human GB1 and produced stereodivergent enzymes with 93% and 79% ee after two rounds.
Problem
Combinatorial protein sequence space is expensive to sample experimentally, limiting efficient exploration of multiple mutations and their interactions.
Method
Machine learning models trained on tested variants screen combinatorial libraries in silico and predict restricted libraries enriched for functional sequences.
Results
Machine learning-assisted evolution reached higher fitness on the GB1 landscape and produced enzymes selective for the S- and R-enantiomers with 93% and 79% ee.
Takeaways & Limitations
In silico modeling enabled larger steps through sequence space and generated diverse sequence solutions for engineering a single parent enzyme toward both product enantiomers.
Takeaways & Limitations
Model performance depends on predictive accuracy, and the approach cannot be expected to outperform a single-mutant walk when predictions are random guesses.
Abstract
from arXiv · showhide
To reduce experimental effort associated with directed protein evolution and to explore the sequence space encoded by mutating multiple positions simultaneously, we incorporate machine learning in the directed evolution workflow. Combinatorial sequence space can be quite expensive to sample experimentally, but machine learning models trained on tested variants provide a fast method for testing sequence space computationally. We validate this approach on a large published empirical fitness landscape for human GB1 binding protein, demonstrating that machine learning-guided directed evolution finds variants with higher fitness than those found by other directed evolution approaches. We then provide an example application in evolving an enzyme to produce each of the two possible product enantiomers (stereodivergence) of a new-to-nature carbene Si-H insertion reaction. The approach predicted libraries enriched in functional enzymes and fixed seven mutations in two rounds of evolution to identify variants for selective catalysis with 93% and 79% ee. By greatly increasing throughput with in silico modeling, machine learning enhances the quality and diversity of sequence solutions for a protein engineering problem.
Results
Machine learning-assisted directed evolution uses combinatorial sequence-function data to predict focused libraries, allowing multiple positions to be explored simultaneously rather than sequentially.
- Results: Machine learning-assisted evolution trains models on sequence and screening data to guide the next library toward higher-fitness variants.The best-performing predicted-library variants become starting points for subsequent evolution rounds.
- Results: The approach samples multiple amino acid residues in each generation and uses combinatorial sequence-function information to restrict later libraries.This increases the probability that experimentally tested variants have high fitness.
- Results: Four positions can be explored simultaneously in one round, enabling broader searches of sequence-function relationships and epistatic interactions.Conventional single-mutation evolution requires N rounds to identify optimal amino acids at N positions.
Validation on an Empirical Fitness Landscape
On the human GB1 empirical fitness landscape, machine learning-assisted evolution more often reached high-fitness solutions than single-mutant walks or recombination under comparable screening burdens, despite imperfect predictive accuracy.
- Validation on an Empirical Fitness Landscape: The simulated single-mutant walk reached 869 fitness peaks, including 533 that outperformed the wild-type sequence.The landscape contained 149,361 reported variants.
- Validation on an Empirical Fitness Landscape: The simulated single-mutation walk fixed the best amino acid at each of four positions sequentially, restricting already explored positions thereafter.This defines the conventional comparison strategy used on the landscape.
- Validation on an Empirical Fitness Landscape: Recombination methods sampled combinatorial libraries and recombined top mutations, with three selected variants allowing a maximum library size of 81.This provided a second conventional comparison strategy.
- Validation on an Empirical Fitness Landscape: For the machine learning comparison, shallow neural networks trained on 470 random variants and tested 100 predictions, matching the single-mutant walk’s screening burden.The study compared the distributions of optimal fitness values found by each approach.
- Validation on an Empirical Fitness Landscape: 8.2% of 600 machine learning simulations reached the global optimum, versus 4.9% for single-mutant walks and 4.0% for recombination runs.All approaches screened the same number of variants.
- Validation on an Empirical Fitness Landscape: Machine learning-assisted evolution had an expected fitness of 6.42, compared with 5.41 for the single-step walk and 5.93 for recombination.The comparison used the same screening burden across approaches.
- Validation on an Empirical Fitness Landscape: The models had low test-set correlation, with Pearson’s r = 0.41 and standard deviation 0.17, yet this accuracy was sufficient to guide evolution.Random-guess predictions would not be expected to outperform a single-mutant walk.
Application to Evolution of Enantiodivergent Enzyme Activity
Machine learning-assisted directed evolution was applied to Rma NOD to generate catalysts selective for either enantiomer of a new-to-nature carbon–silicon bond-forming reaction. Two rounds of restricted-library prediction explored seven positions and produced selective enzymes while concentrating experiments in sequence-space regions enriched in function.
- Reaction and starting enzyme: Rma NOD catalyzes a new-to-nature carbon–silicon bond-forming reaction, with the Y32K/V97L starting variant showing 76% ee for the (S)-enantiomer.The reaction uses phenyldimethyl silane and Me-EDA in whole-cell reactions.
- Library design: Set I randomized positions 32, 46, 56, and 97, while Set II randomized positions 49, 51, and 53 after beneficial Set I mutations were fixed.Random input libraries were used to train models, followed by restricted predicted libraries enriched for high selectivity.
- Evolution outcome: After two rounds, variants with 93% and 79% ee were obtained for the (S)- and (R)-enantiomers, respectively.The workflow used 445 sequence–function relationships for training and tested 360 additional predicted variants, totaling 805 experimentally tested variants across seven positions.
- Diversity and activity: The approach identified multiple sequence solutions and increased cellular activity, which can improve apparent selectivity by overcoming nonselective background activity.The two most selective (S)-variants differed by less than 1% ee, while one had higher total activity.
- Scope: The authors caution that other evolution strategies may sometimes discover higher-fitness variants more quickly, so library creation should not be judged from isolated examples alone.The stated purpose of the approach is to increase the likelihood of success rather than guarantee the best variant in every instance.
- Sequence-space enrichment: Machine-learning predictions shifted tested libraries toward regions of sequence space enriched in high fitness for both enantiomers.In Set I, 90 predicted variants were tested from libraries of 864 and 630 variants; in Set II, 90 were tested from libraries of 192 and 90 variants for the (S)- and (R)-enantiomers, respectively.
Discussion
Machine learning-assisted directed evolution uses combinatorial sequence sampling to reduce experimental screening while identifying diverse, high-fitness protein variants. The approach models sequence-function relationships and can navigate sequence space through simultaneous mutations.
- Machine learning rapidly screens combinatorial libraries in silico, reducing in vitro screening while identifying variants for selective catalysis.The study evolved a single parent enzyme toward both product enantiomers of a new-to-nature C–Si bond-forming reaction.
- Simultaneous incorporation of multiple mutations accelerates evolution by navigating different fitness-landscape regions and avoiding low-fitness search endpoints.This strategy can identify combinations of beneficial mutations and sequences unavailable by independently recombining the best amino acids at each position.
- Distinct sequence solutions can produce similar properties: tyrosine and arginine at position 49 retain enantioselectivity despite proline being conserved in two highly selective variants.The passage reports less than 1% loss in enantioselectivity for tyrosine or arginine at that position.
- Machine learning can leverage limited experimental resources without requiring a detailed mechanistic understanding of protein function.The models provide efficient estimates of desired properties for proteins in large libraries.
- The empirical GB1 validation used measured fitness values for 149,361 of 160,000 sequences across four positions and compared three simulated evolution strategies.The strategies were single-mutation walks, simulated recombination, and machine learning-assisted directed evolution.
I. General procedures
The study combines machine-learning models with experimental directed evolution to prioritize combinatorial protein variants while reducing experimental screening demands. It applies this workflow to empirical fitness-landscape simulations and enzyme engineering, using model predictions, degenerate-codon libraries, and experimental validation.
- Model training: Machine-learning models were trained on sequencing and enantiomeric data using diverse regressors, including linear, tree-based, ensemble, and neural-network methods.The model panel included K-nearest neighbors, linear regressors, decision trees, random forests, and multilayer perceptrons.
- Library design: Predicted libraries were encoded with degenerate codons by using amino-acid frequencies from top-ranked model predictions, while retaining amino-acid combinations identified as optimal by the models.All nine models were considered when selecting amino acids, although this procedure can generate non-optimal combinations.
- Experimental validation: After two rounds of machine-learning-assisted evolution, the study evaluated selective variants and compared their activity with the starting sequence.Tables summarize modeling statistics, enantioselectivity, and relative activity for input and predicted libraries.
- Practical limitations: A practical limitation is that cloning by degenerate codons loses the exact sequence combinations predicted by the models and may introduce non-optimal combinations.Direct synthesis of top variants would avoid this interpretation step but is described as expensive.
- Fitness-landscape simulations: 570 variants are tested in all approaches, enabling comparison of machine-learning-assisted evolution with directed-evolution simulations at matched screening burden.The comparison included iterative site-saturation mutagenesis and recombination runs on the GB1 empirical landscape.
- Fitness-landscape simulations: 8.4% of machine-learning-assisted simulations reached the global optimum, compared with 4.9% of all starting sequences.The supplementary analysis reports similar performance to directed evolution with 300–400 total tested sequences.
V. Library coverage
Library coverage depends strongly on how many variants are tested and how codons distribute amino acids. The comparison also shows that machine-learning-assisted evolution can match single-mutation walks with fewer tested variants in this case.
- Coverage assumptions: 95% library coverage is often approximated as three times the library size, so 19 mutations require 57 tested variants under the stated rule.This estimate assumes equal amino-acid frequencies.
- Codon choice: 944 variants are required for 10 NNS/NNK libraries because methionine occurs at frequency 1/32.The codon choice changes the least frequent amino-acid frequency and therefore the number needed for 95% coverage.
- Codon choice: 644 variants are required for 10 libraries using NDT/VHG/TGG codons, known as the 22c-trick, with methionine occurring 1/22 times.This provides a balance between degenerate-codon complexity and amino-acid coverage.
- Experimental comparison: The main-text directed-evolution baseline uses 570 variants, while the single-mutation walk performs similarly to machine learning at 300–400 variants.The 570-variant baseline comes from oversampling 10 libraries containing 19 desired variants each.