Source-linked AI summary

General Multimodal Protein Design Enables DNA-Encoding of Chemistry

Jarrid Rector-Brooks, Théophile Lambert, Marta Skreta, Daniel Roth, Yueming Long, Zi-Qi Li, Xi Zhang, Miruna Cretu, Francesca-Zhoufan Li, Tanvi Ganapathy, Emily Jin, Avishek Joey Bose, Jason Yang, Kirill Neklyudov, Yoshua Bengio, Alexander Tong, Frances H. Arnold, Cheng-Hao Liu

arXiv:2604.05181v1cs.LG

TL;DR

New-to-nature enzyme design lacks obvious starting proteins and existing methods rely on predefined catalytic motifs. DISCO jointly designs protein sequence and structure around arbitrary biomolecules, using multimodal inference-time steering. Conditioned only on reactive intermediates, it generated diverse, evolvable carbene transferases with novel active sites and high activities across four new-to-nature reactions.

  • Problem

    De novo enzyme design for new-to-nature reactions remains limited because existing methods rely on predefined active-site geometries and sequential backbone-to-sequence generation.

  • Method

    DISCO jointly designs protein sequences and 3D structures around arbitrary biomolecules and optimizes both modalities with inference-time steering.

  • Results

    Conditioned solely on reaction intermediates, DISCO generated diverse carbene transferases with novel active sites and activities surpassing typical directed-evolution starting points and, in selected cases, evolved variants.

  • Takeaways & Limitations

    DISCO provides a global search for diverse, evolvable enzymes and expands the searchable space of DNA-encoded chemical reactivity.

  • Takeaways & Limitations

    The reactions explored represent only a small subset of synthetically valuable transformations absent from nature, and more complex mechanisms may require tighter integration with directed evolution and richer biophysical constraints.

Abstract

from arXiv · show

Evolution is an extraordinary engine for enzymatic diversity, yet the chemistry it has explored remains a narrow slice of what DNA can encode. Deep generative models can design new proteins that bind ligands, but none have created enzymes without pre-specifying catalytic residues. We introduce DISCO (DIffusion for Sequence-structure CO-design), a multimodal model that co-designs protein sequence and 3D structure around arbitrary biomolecules, as well as inference-time scaling methods that optimize objectives across both modalities. Conditioned solely on reactive intermediates, DISCO designs diverse heme enzymes with novel active-site geometries. These enzymes catalyze new-to-nature carbene-transfer reactions, including alkene cyclopropanation, spirocyclopropanation, B-H, and C(sp$^3$)-H insertions, with high activities exceeding those of engineered enzymes. Random mutagenesis of a selected design further confirmed that enzyme activity can be improved through directed evolution. By providing a scalable route to evolvable enzymes, DISCO broadens the potential scope of genetically encodable transformations. Code is available at https://github.com/DISCO-design/DISCO.

Introduction

New-to-nature enzyme design remains largely unrealized because existing approaches require known active-site geometries and sequential backbone-then-sequence generation. DISCO addresses this gap by jointly designing sequences and structures around arbitrary biomolecules without predefined residue motifs.

  • Existing generative models create novel binders and structural motifs, but de novo enzyme design for new-to-nature reactions remains largely unrealized.
  • Current pipelines depend on pre-specified active-site residue arrangements or theozymes, excluding reactions without known motifs or mechanisms.
  • Sequential backbone generation followed by inverse folding limits joint adaptation of protein sequence and structure during design.
  • DISCO simultaneously designs protein sequences and 3D structures de novo while conditioning on arbitrary biomolecules without predefined residue motifs.

Multimodal protein generation with DISCO

DISCO jointly generates protein sequences and 3D structures, conditions on diverse biomolecular contexts, and uses inference-time steering to optimize multimodal design objectives. Its designs are foldable, diverse, chemically responsive, and often novel relative to known proteins and motifs.

  • Multimodal protein generation with DISCO: DISCO models sequences and 3D structures as a joint distribution denoised through a unified generative process.Masked discrete diffusion handles sequences, while continuous diffusion handles 3D atomic coordinates.
  • Multimodal protein generation with DISCO: Cross-modal recycling lets sequence predictions use emerging structural features while structural predictions adapt to evolving sequence identity.The mechanism conditions each generation step on multiple sequence and structure encodings.
  • Multimodal protein generation with DISCO: Approximately 90% of unconditional monomer sequences refold within 2 Å RMSD of their designed backbones using ESMFold.Self-correction, sequence temperature, and noisy guidance substantially improve sequence-structure co-designability.
  • Conditioning on arbitrary biomolecular contexts: DISCO conditions on small molecules, metallocofactors, reactive intermediates, nucleic acids, and multiple ligands that co-fold with the designed protein.The STUDIO-179 benchmark spans 179 natural and non-natural ligands across catalysis, pharmaceuticals, luminescence, and sensing.
  • Inference-time steering: FKC-MM jointly steers sequence and structure toward multimodal rewards, while FKC-SG promotes target binding and penalizes binding to structurally similar decoys.These methods steer generation directly rather than relying only on generate-and-filter strategies.
  • DISCO designs exhibit realistic protein features with novel, complementary motifs: Joint co-design produces chemically responsive pockets, including ligand-matched lipophilicity, cofactor-coordinating residues, sufficiently large cavities, and valid alternative ligand conformers.The designs also exhibit natural-like protein statistics, diverse topologies, and favorable geometric and surface properties.
  • DISCO designs exhibit realistic protein features with novel, complementary motifs: Most binding motifs in co-designable generations lack natural homologs in AlphaFoldDB, while generated motifs show over 90% cluster diversity.Novelty is reported alongside foldability and physicochemical realism.

Designing enzymes for new-to-nature biocatalysis

DISCO designs diverse carbene-transfer enzymes by conditioning sequence–structure generation on reactive intermediates rather than fixed catalytic motifs. Across four new-to-nature reactions, selected designs showed substantial activity, novel structural solutions, and evidence of evolvability.

  • Design strategy: Conditioning on heme–carbene intermediates lets DISCO explore catalytic solutions without theozyme scaffolding or fixed residue assumptions.The heme–carbene intermediate is a key rate-determining species in these reactions.
  • Design strategy: 90 experimentally tested designs formed 75 distinct structural clusters, with no pair sharing more than 50% sequence identity.The designs were selected from approximately 10^4 generated sequence–structure pairs after computational filtering.
  • Catalytic scope: DISCO designs catalyzed cyclopropanation, B–H insertion, C(sp3)–H insertion, and spirocyclopropanation across chemically distinct substrates.All 90 designs were screened in whole-cell Escherichia coli format across the four reactions.
  • Catalytic scope: 98% yield and 5,170 TTN were achieved for B–H insertion, while C(sp3)–H insertion reached 42% yield and 2,360 TTN.The B–H result exceeded the prior starting point and laboratory-evolved variant; the C(sp3)–H result rivaled a previously evolved P411-CHF catalyst.
  • Evolvability: A single error-prone-PCR round produced approximately 35 dCT-H11 variants with improved activity and divergent enantioselectivity.Screening covered approximately 700 mutants, with enantiomeric preference shifting from +35% to +49% ee or inverting to +35% to -35% ee.
  • Structural novelty: Selected designs used structurally novel active-site arrangements and unrelated protein folds, including dCT-H11’s 21% identity to a noncatalytic TetR-family match.The closest active-site motif comparison for dCT-H11 showed RMSD > 7 Å, while other designs had no corresponding AlphaFoldDB motif.

General biomolecular design unlocks new-to-nature reactivities

DISCO couples protein sequence and 3D-structure generation while conditioning on arbitrary biomolecules, enabling catalytic design without predefined transition-state or residue motifs. From 90 genes, its designs catalyzed four new-to-nature reactions and produced active-site residues with evidence of subsequent evolvability, although the explored reactions remain a small subset of useful chemistry.

  • General biomolecular design: DISCO generates protein sequences and 3D structures as a coupled object and can condition them on arbitrary biomolecules.The framework uses inference-time steering across both modalities to make generative sampling a controllable search process.
  • New-to-nature reactivities: Conditioning solely on reaction intermediates produced promiscuous carbene transferases across four new-to-nature reactions from 90 genes.This bypassed transition-state calculations, theozyme scaffolding, and extensive wet-lab screening.
  • New-to-nature reactivities: Top designs achieved 98% yield for B–H insertion and more than 2,300 TTN for C(sp3)–H insertion, exceeding previously evolved biocatalysts.These results establish activity across multiple new-to-nature carbene-transfer transformations.
  • Scope and outlook: The demonstrated reactions represent only a small subset of synthetically valuable transformations absent from nature.The paper identifies closing the loop between generative discovery and directed evolution as important for increasingly complex mechanisms.

Supplementary Information

The supplementary framework models protein sequence and structure with separate diffusion processes before combining them into DISCO’s joint generative process. Discrete sequence diffusion uses masking, while inference can plan token unmasking and remasking across both modalities.

  • Continuous-space diffusion modeling: Continuous diffusion gradually destroys structural information with an SDE and learns a score-based reverse process for generation.The forward process is designed so the terminal distribution is approximately standard normal.
  • Discrete-space diffusion modeling: Discrete diffusion masks sequence tokens and trains a denoiser to predict token distributions across masking levels.The training objective is a weighted cross-entropy loss over the masking schedule.
  • Discrete-space diffusion modeling: Standard discrete denoising randomly selects masked positions and permanently commits unmasked tokens, limiting generation-order control and error correction.These limitations motivate alternative reverse-time processes.
  • Discrete-space diffusion modeling: P2 addresses both discrete-diffusion limitations by planning which positions to unmask and remasking already unmasked tokens for possible revision.In practice, the planners are uniform because confidence-dependent remasking reduces structural diversity.
  • Multi-modal diffusion modeling: DISCO combines sequence and structure diffusion into a joint reverse process driven by a single denoiser over both modalities.Inference follows trajectories of sequence and structure noise levels from fully noisy to clean states.

A.3 Data pipeline

The data pipeline trains DISCO on a weighted, clustered protein dataset with several cropping strategies, while the architecture adapts AlphaFold 3 for joint sequence–structure co-design. Its denoiser repeatedly exchanges information between modalities and outputs both sequence and structure estimates.

  • Data pipeline: DISCO trains on a weighted PDB dataset processed with AlphaFold 3 clustering and filtering procedures.The dataset includes single chains and chain-pair interfaces, with interfaces defined by a minimum heavy-atom separation below 5 Å.
  • Data pipeline: Samples use contiguous, spatial, and spatial-interface crops with weights of 20%, 40%, and 40%, respectively.Cropping is performed per sample.
  • Model architecture: The AlphaFold 3-based architecture removes template and MSA modules, replacing MSA-derived evolutionary information with the DPLM 650M protein language model.These changes address the computational difficulty of evolving sequences during inference.
  • Denoiser architecture: The denoiser iteratively refines single and pair representations through four cycles, using cross-modal sequence–structure updates and producing denoised sequence and structure estimates.Pair representations also incorporate relative positions, bond features, and protein-language-model information.

A.5 Training

DISCO is trained with joint sequence-and-structure diffusion losses and selected using a tradeoff between co-designability and structural diversity. The final checkpoint and inference procedures reflect this multi-objective selection.

  • Training setup: DISCO contains 888 million parameters, including 235 million trainable parameters, and was trained for 160,000 steps across 32 L40S GPUs for 11 days.Unresolved or unreliable atomic positions and residues are masked when evaluating losses.
  • Joint diffusion architecture: The diffusion module processes noisy sequences and structures jointly through conditioning, cross-modal encoding, atom attention, and a diffusion transformer.The module predicts sequence logits and denoised structural coordinates through modality-specific and shared representations.
  • Training objectives: Structure training uses weighted aligned MSE, Smooth LDDT, and distogram losses, while sequence training uses masked diffusion loss.The full objective combines sequence and structure losses with weights αseq = 1, αMSE = 4, αsmooth_lddt = 4, and αdistogram = 0.03.
  • Training objectives: Unreliable structural data are excluded from training losses, including unresolved atoms and residues with insufficient backbone-atom occupancy.The structure diffusion procedure also applies optimal ground-truth chain assignment before computing losses.
  • Model selection: The final checkpoint was selected by evaluating EMA checkpoints for the best tradeoff between co-designability and structural cluster designability.Checkpoint 160k was selected because it provided superior tradeoffs, although structural diversity later decreased after peaking around 200k steps.
  • Inference: Inference settings changed the same checkpoint’s co-designability from 16% to 88%, showing that sampling choices strongly affect generation quality.The reported settings improved co-designability while structural diversity and co-designability could trade off against one another.

A.7.2 FKC-SG experimental details

This section describes FKC-SG experiments for steering DISCO toward target-specific binding while avoiding structurally related decoys, alongside multimodal reward-guidance details and evaluation procedures.

  • Experimental setup: FKC-SG was evaluated on valine–proline, aldosterone–cortisone, and biotin–PLP on/off-target pairs across inverse temperatures β ∈ {−0.5, −1.0, −2.0}.For each ligand pair, the first molecule was the on-target and the second the off-target.
  • Experimental setup: 200K protein sequence–structure samples were generated per condition for 150-residue proteins, then filtered by on-target co-designability before off-target discrimination was assessed.A design passed the initial filter when both ligand-centroid and protein RMSDs were below 2 Å after on-target refolding.
  • Multimodal rewards: Multimodal guidance rewards sequence–structure coupling through soft residue-contact counts, including disulfide bonds and cation-π interactions.The disulfide reward uses Cβ distances, while the cation-π reward distinguishes cationic Arg/Lys from aromatic Phe/Tyr/Trp/His residues.

A.8.4 Conditional evaluation of small-molecule ligands

This section defines a 179-molecule benchmark for conditional protein–ligand generation and specifies the structure-prediction, co-designability, clash, and binding-site analyses used for evaluation.

  • Benchmark: STUDIO-179 contains 179 chemically diverse molecules spanning cofactors, metals, drug-like compounds, fluorophores, saccharides, pollutants, and metabolites.The benchmark includes multiple different molecules within these categories.
  • Generation protocol: Fifty protein–ligand complexes were generated for each of three target lengths—150, 200, and 250 residues—for every benchmark molecule.Molecule geometries were optimized with GFN2-xTB and supplied to DISCO with bond and reference-conformer information.
  • Evaluation criteria: A design was co-designable when refolding satisfied RMSDbackbone(Xdesign, XChai) < 2.0 Å and the corresponding ligand-centroid criterion.For multi-ligand systems, the ligand criterion had to hold independently for every ligand.
  • Evaluation criteria: Protein–ligand steric clashes were flagged when any protein and ligand heavy atoms were closer than the van der Waals radius threshold with a 0.5 Å tolerance.The criterion was applied to backbone protein structures and ligand heavy atoms.
  • Binding-site analysis: Binding-site analyses defined the site as residues whose backbone atoms lie within 10 Å of ligand or cofactor heavy atoms, retaining up to 10 closest residues.This definition was used for logP and pocket-hydrophobicity analyses.

A.8.5 Conditional evaluation: nucleic acid binding

This section details conditional evaluation for nucleic-acid binding and the broader design, filtering, structural-novelty, and computational chemistry procedures used in the study.

  • Nucleic-acid evaluation: DNA and RNA targets were evaluated using 100 generated samples per target, with protein lengths sampled uniformly from 50 to 80 residues.DNA chains were cropped to a maximum length of 10, and complexes were refolded with Chai-1.
  • Nucleic-acid evaluation: Nucleic-acid complexes were co-designable when the protein backbone RMSD after Chai-1 refolding was below 2.0 Å while preserving the designed protein–nucleic-acid relationship.Nucleotide reference atoms were aligned using phosphorus atoms as the primary anchor.
  • Baselines: DISCO was compared with RFdiffusion3 and BoltzGen for conditional small-molecule and nucleic-acid generation, with RFDiffusion All-Atom additionally used for small molecules.Other unconditional baselines were excluded from conditional tasks because they did not natively support ligand or nucleic-acid conditioning.
  • Motif analysis: Binding-motif novelty was assessed with Folddisco, classifying motifs as novel when unmatched or when a full match exceeded 3.0 Å RMSD.Motifs were defined from residues closest to the ligand in Chai-1-refolded structures.
  • Motif analysis: Motif diversity used nearest-neighbour metrics including Cα-RMSD, chemical cost, 1 − lDDT, Frobenius distance, and cluster ratio.A cluster ratio near 1 indicates nearly every design has a structurally distinct active-site configuration, whereas a ratio near 0 indicates collapse onto shared geometries.
  • Enzyme design campaigns: Enzyme campaigns generated approximately 10^4 sequence–structure pairs conditioned on reactive intermediates, followed by Chai-1 and AlphaFold 3 co-designability filters and active-site physicochemical checks.Retained designs had to pass both structure-prediction oracles, with filters covering contacts, burial, and metal coordination.

A.11.1 Additional unconditional generation results

Additional unconditional-generation analyses examine co-designability, diversity, and sample properties, emphasizing that quality metrics should be considered alongside folding success rather than optimized in isolation.

  • Unconditional generation: Co-designability is very high for unconditional generation, and filtering samples for co-designability does not qualitatively change their analyzed protein properties.The analyses include amino-acid composition, secondary structure, radius of gyration, and relative contact order.
  • Evaluation metrics: Table S4 identifies co-designable structural and sequence diversity as informative metrics because they jointly reflect sample quality and absence of mode collapse.The table cautions that maximizing co-designability alone is inappropriate because folding models cannot refold every PDB protein within 2 Å RMSD.
  • Length dependence: Figure S5 reports unconditional co-designability distributions and the proportion of co-designable clusters across protein lengths.These plots examine how co-designability and cluster proportions vary with chain length.
  • Sample properties: Figures S6–S8 compare generated amino-acid distributions, long-range contact types, and radii of gyration with sampled PDB training structures.These analyses assess whether unconditional samples reproduce broad structural and sequence properties of the training distribution.
  • Disulfide content: Figure S9 characterizes disulfide-bond counts in PDB protein chains of length at most 100.This distribution provides a reference for short-chain disulfide content.

A.11.2 Additional conditional generation results

Additional analyses characterize DISCO’s conditional designs across natural and non-natural ligands, showing broad co-designability and chemically structured protein–ligand interfaces. The generated proteins exhibit diverse structural and chemical properties, although co-designability filtering can introduce bias.

  • Filtering effects: Co-designability filtering produced generally similar qualitative protein analyses, with a significant change primarily observed in secondary-structure distributions.The authors evaluate most general properties without co-designability filtering, reserving filtering mainly for motif novelty and diversity analyses.
  • Benchmark performance: DISCO generated the highest proportion of diverse, co-designable sequence–structure pairs across all tested ligands.The benchmark covered the remaining STUDIO-179 ligands, ranging from natural to non-natural molecules.
  • Generated-protein properties: Conditional designs were analyzed across protein length, amino-acid composition, secondary structure, contact order, surface properties, backbone geometry, and radius of gyration.These analyses were compared with 50,000 randomly sampled PDB structures from the training data.
  • Design examples: Additional examples include luciferin and a substituted naphthalene diimide, demonstrating conditional designs for both natural and non-natural ligands.The examples use distinct protein topologies: a hybrid helix-and-sheet topology and a helix bundle.
  • Motif generation: Generated motifs include a class I type II Cu center and novel residue motifs with no Folddisco match in many cases.The Cu motif uses 2 His, 2 Cys, and 1 Glu residues in tetrahedral coordination with Cu[2+].

B.4.2 Analysis by LC–MS (compound 4b for directed evolution)

This section documents analytical procedures and validation data for carbene-transfer products, including chromatographic analyses, enantioselectivity, diastereoselectivity, and directed-evolution variants. It also records the mutations introduced into evolved designs.

  • Analytical methods: HPLC–MS, chiral HPLC–UV, and chiral GC-FID were used to analyze products and directed-evolution samples.Compound 4b evolution samples were analyzed by chiral GC-FID, while products 1b–3b were analyzed by chiral HPLC–UV.
  • Enantioselectivity: 61:39, 63:37, and 40:60 enantiomeric ratios were reported for dCT-H11 or dCT-F9 across cyclopropanation, B–H insertion, and C–H insertion products.The ratios correspond respectively to dCT-H11 for 1b, dCT-F9 for 2b, and dCT-H11 for 3b.
  • Selectivity validation: 99:1 trans:cis diastereomeric ratio was measured for dCT-H11 in selected methoxystyrene cyclopropanation validation results.The corresponding product is compound 1b.
  • Directed evolution: Error-prone PCR generated evolved variants containing multiple amino-acid substitutions relative to the parent design.Variants included dCT-r1-B7, dCT-r1-D1, dCT-r1-E3, dCT-r1-E12, and dCT-r1-F1.
  • Constructs: The selected protein constructs included dCT-H11, dCT-H10, dCT-F9, dCT-G9, and dCT-G1 sequences used in validation experiments.The supplied records list their corresponding DNA sequences.

B.9 Calibration curves for GC-FID analysis

Calibration curves were prepared for GC-FID analysis of compounds 1b, 2b, and 4b using internal standards and multiple analyte-to-standard concentrations. Separate preparation procedures were used for compound 2b.

  • Calibration design: Seven standard-to-internal-standard ratios were prepared for compounds 1b, 2b, and 4b calibration curves.The reported ratios included 1:1, 0.5:1, 0.25:1, 0.125:1, and 0.0625:1.
  • Internal standards: 1b, 4b, and related analytes used 1,2-diphenylethane as the internal standard in cyclohexane:EtOAc.Compound 2b instead used 1,2,3-trimethoxystyrene as the internal standard.

B.10 Calibration curves for LC-MS analysis

This section records LC–MS calibration and supporting analytical traces for compounds used in synthesis and enzyme assays. It includes compound 4a calibration, compound 4b LC–MS calibration, product chromatograms, NMR characterization, and chiral GC-FID traces.

  • LC–MS calibration: Compound 4a calibration used successive concentrations from 10 to 300 µM with papaverine hydrochloride as the internal standard.The listed concentrations were 10, 30, 50, 100, 200, and 300 µM.
  • Chiral analysis: Chiral GC-FID traces compare racemic and enantiopure 4b standards with reactions catalyzed by dCT-H11 and evolved variants.The evolved variants include dCT-r1-B7, dCT-r1-E3, dCT-r1-E12, dCT-r1-F1, and dCT-r1-D1.
Loading 2604.05181v1…