Source-linked AI summary
Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute
Kieran Didi, Zuobai Zhang, Guoqing Zhou, Danny Reidenbach, Zhonglin Cao, Sooyoung Cha, Tomas Geffner, Christian Dallago, Jian Tang, Michael M. Bronstein, Martin Steinegger, Emine Kucukbenli, Arash Vahdat, Karsten Kreis
TL;DR
Structure-based binder design has split between conditional generation and structure-predictor-guided hallucination, while paired binder-target data remains limited. Proteina-Complexa combines a flow-based generative prior, Teddymer synthetic complexes, and inference-time optimization. It reports state-of-the-art in-silico performance across binder, small-molecule, and enzyme-design tasks, including stronger results than prior hallucination methods under normalized compute budgets.
Problem
Structure-based binder design is divided between generative modeling and hallucination, and expressive binder generators lack abundant paired binder-target multimer data.
Method
Proteina-Complexa extends atomistic flow-based generation with Teddymer synthetic complexes and uses the resulting generative prior for inference-time optimization.
Results
Proteina-Complexa achieves state-of-the-art in-silico binder-design performance across protein and small-molecule targets, enzyme design, and normalized-compute comparisons with hallucination methods.
Takeaways & Limitations
The framework unifies generation and optimization without sequence re-design while supporting hydrogen-bond optimization and fold-class-guided binder diversity.
Takeaways & Limitations
Applying Monte Carlo Tree Search to the flow model’s continuous state/action space requires further technical innovations.
Abstract
from arXiv · showhide
Protein interaction modeling is central to protein design, which has been transformed by machine learning with applications in drug discovery and beyond. In this landscape, structure-based de novo binder design is cast as either conditional generative modeling or sequence optimization via structure predictors ("hallucination"). We argue that this is a false dichotomy and propose Proteina-Complexa, a novel fully atomistic binder generation method unifying both paradigms. We extend recent flow-based latent protein generation architectures and leverage the domain-domain interactions of monomeric computationally predicted protein structures to construct Teddymer, a new large-scale dataset of synthetic binder-target pairs for pretraining. Combined with high-quality experimental multimers, this enables training a strong base model. We then perform inference-time optimization with this generative prior, unifying the strengths of previously distinct generative and hallucination methods. Proteina-Complexa sets a new state of the art in computational binder design benchmarks: it delivers markedly higher in-silico success rates than existing generative approaches, and our novel test-time optimization strategies greatly outperform previous hallucination methods under normalized compute budgets. We also demonstrate interface hydrogen bond optimization, fold class-guided binder generation, and extensions to small molecule targets and enzyme design tasks, again surpassing prior methods. Code, models and new data will be publicly released.
1 INTRODUCTION
Proteina-Complexa unifies conditional generative modeling with inference-time optimization for fully atomistic binder design. It combines a pretrained generative prior, synthetic and experimental multimer data, and compute-scaled search to improve in-silico design across multiple targets and tasks.
- Structure-based binder design has followed conditional generation or sequence optimization using structure-predictor scores, but Proteina-Complexa treats these as complementary rather than exclusive.
- Teddymer supplies large-scale synthetic binder-target complexes assembled from domain-domain interactions in predicted monomer structures, addressing limited experimental multimer data.The dataset is combined with diverse monomers and experimental multimer structures during staged training.
- Complexa extends La-Proteina with latent target conditioning and uses its generative prior to guide best-of-N sampling, beam search, Feynman–Kac steering, and Monte Carlo Tree Search.Structure predictors provide interface confidence scores or hydrogen-bond energies as inference-time rewards.
- Complexa outperforms prior generative models on in-silico binding metrics and prior hallucination methods under normalized compute budgets across protein and small-molecule binder tasks.The framework also reports a large-margin improvement on an enzyme-design benchmark without sequence re-design.
- The framework additionally supports interface hydrogen-bond optimization and fold-class conditioning for controllable binder diversity.
- Complexa is presented as a structure-based protein-design method that scales both training data and inference compute, with code, model weights, and Teddymer planned for public release.
2 BACKGROUND AND RELATED WORK
The paper builds on flow-matching protein generation and situates binder design between generative modeling and hallucination. Its related-work context emphasizes atomistic generation, complementary prior work, and limited paired binder-target data.
- La-Proteina uses partially latent flow matching to generate fully atomistic monomers and analyze their biophysical validity and motif scaffolding.
- The present work focuses specifically on protein binder design, making it orthogonal and complementary to La-Proteina’s monomer-generation work.
- Generative binder-design methods train flow or diffusion models on binder-target complexes, whereas hallucination methods optimize sequences using structure-predictor feedback.
3 PROTE´INA-COMPLEXA
Complexa combines a latent, fully atomistic binder generator with inference-time search and optimization, using target conditioning, synthetic dimer data, and structure-prediction rewards. Its design supports protein and small-molecule targets, interface hydrogen-bond optimization, and fold-class conditioning.
- Training data: Teddymer supplies synthetic binder-target dimers derived from domain-domain interactions, complementing limited experimental multimer data.The training pipeline combines diverse monomers, Teddymer dimers, and experimental multimer structures.
- Base generative model: Complexa extends La-Proteína with latent target conditioning while keeping the autoencoder focused on monomeric binders.Target features are concatenated with binder coordinates and latent variables for joint processing by the flow-model denoiser.
- Target conditioning: Complexa represents small-molecule targets with atomic features, molecular bond information, and distances to binder backbone atoms.These target features are embedded and concatenated with binder embeddings, while pair features encode internal bonds and binder-target distances.
- Training objectives: Translation noise is added to binder alpha-carbon coordinates to force reasoning about global interface positioning.The method uses random global translation noise during interpolation; the paper identifies this as critical for binder placement.
- Evaluation and rewards: Structure-prediction confidence and alignment scores serve both as binder-quality metrics and as rewards for inference-time optimization.The framework also optimizes interface hydrogen-bond energies on refolded generated binder sequences.
- Inference-time optimization: Complexa adapts beam search and Feynman–Kac steering to search within the generative prior during binder generation.Feynman–Kac steering samples from a reward-tilted model distribution, while the broader framework includes test-time scaling methods from diffusion models.
4 EXPERIMENTS
Experiments show that Complexa performs strongly across protein and small-molecule binder design, while inference-time search improves success under matched compute. The framework also supports interface hydrogen-bond optimization, difficult targets, fold-guided generation, and enzyme design.
- Generative performance: Complexa’s generative model is evaluated against RFDiffusion-AllAtom on four small-molecule targets without inference-time optimization.The comparison uses Complexa’s self-generated sequences, while RFDiffusion-AllAtom uses LigandMPNN.
- Inference-time scaling: Across easy and hard protein targets, Complexa’s inference-time methods lead overall under matched compute, although BindCraft is ahead on some targets.Best-of-N is sufficient for easy targets, whereas Beam Search, FKS, and MCTS are needed on hard targets.
- Inference-time scaling: Initializing BindCraft from Complexa samples accelerates search on easy targets but not hard targets.The hard-target VEGFA case study highlights superior performance from Complexa’s own inference-time optimization methods without sequence redesign.
- Interface optimization: Optimizing interface hydrogen-bond reward alongside fipAE increases average unique success rate and substantially increases interface hydrogen bonds.The comparison is reported for beam search with different folding and hydrogen-bond reward combinations.
- Challenging targets: With extended searches beyond 100 GPU hours, Complexa finds 15 unique TNF-α successes, 7 H1 successes, and 1 IL17A success.These challenging multi-chain targets produced no successes from publicly available baselines within less than 32 GPU hours.
- Enzyme design: Complexa significantly outperforms RFDiffusion2 in 38/41 AME enzyme-design benchmark tasks.The result holds with both redesigned and self-generated sequences.
5 CONCLUSIONS
The conclusion presents Complexa as a fully atomistic framework combining generative modeling with inference-time optimization. It reports state-of-the-art in-silico binder design and highlights hydrogen-bond optimization, while noting broader deployment risks and planned release of resources.
- Contributions: Complexa bridges large-scale generative modeling with test-time compute scaling for fully atomistic protein binder generation.The framework is pretrained on Teddymer, a synthetic dimer dataset built from AFDB domain–domain interactions.
- Contributions: Complexa achieves state-of-the-art de novo binder design without sequence redesign and outperforms prior hallucination methods by unifying generation and optimization.The conclusion frames this unification as combining previously separate generative and hallucination approaches.
- Extensions: Interface hydrogen-bond optimization demonstrates the framework’s flexibility and supports integrating physics- and learning-based modeling.The conclusion identifies this as an opportunity enabled by the framework rather than a completed broader integration.
- Future work: The authors restrict current applications to protein and small-molecule targets and identify other molecular modalities as future extensions.The stated future direction includes a unified model for proteins, peptides, small molecules, nucleic acids, antibodies, and related modalities.
- Responsible use: Generative binder-design models could support medicine, biotechnology, and basic science but require prudent oversight because they may be misapplied.The ethics statement specifically emphasizes responsible deployment.
A LIMITATIONS AND FUTURE WORK
The supplied material describes La-Proteína’s atomistic latent architecture and flow-matching generation process. It also indicates that future work should extend evaluation beyond current molecular targets and computational settings.
- Representation: La-Proteína represents alpha-carbon coordinates explicitly while encoding sequence and other atomistic details in fixed-size per-residue latent variables.A VAE maps between full protein structures and this partially latent representation.
- Autoencoder: The VAE decoder takes latent variables and alpha-carbon coordinates as input and outputs distributions over sequences and non-alpha-carbon coordinates.The sequence distribution is categorical, while the non-alpha-carbon coordinate distribution is Gaussian.
- Autoencoder: The encoder maps complete proteins to latent representations, parameterizing a factorized Gaussian over latent variables.Its inputs include alpha-carbon coordinates, other atom coordinates, and sequence.
- Generation: Flow matching learns the joint distribution of alpha-carbon coordinates and latent variables in continuous fixed-size space.Sampling transports standard Gaussian noise toward the target data distribution.
- Generation: La-Proteína generates samples by numerically simulating stochastic differential equations from initial to final interpolation times.The equations use a score derived from the learned velocity field.
- Sampling: Low-temperature noise scaling and distinct coordinate schedules improve co-designability during sampling.Noise parameters are typically set below one, with alpha-carbon coordinates denoised faster than latent variables.
B.5 ARCHITECTURES
The architecture combines transformer-based atomistic generation with a large synthetic dimer dataset derived from AFDB and TED annotations. Teddymer is filtered and clustered before training uses its representatives.
- Architecture: La-Proteína’s encoder, decoder, and denoiser use transformers with pair-biased attention mechanisms.The denoiser additionally conditions on interpolation times through adaptive layer normalization and output scaling.
- Dataset construction: Teddymer processing begins by treating TED-annotated AFDB domains as chains and retaining structures from the AFDB50 clustered database.AFDB contains about 203 million structures, motivating this initial filtering step.
- Dataset construction: The dimer database contains 123,606,001 extracted dimers, filtered using proximity and CATH annotation criteria.A dimer requires each chain to have at least four residues within 10 Å of the other chain.
- Dataset construction: Structural and interface clustering produces 3,556,223 dimer clusters.Clustering uses GPU-accelerated Foldseek-Multimer with chain-level structural and interface similarity.
- Dataset construction: Final Teddymer filtering yields 510,454 cluster representatives and 7,112,609 overall datapoints, with training performed only on representatives.The filters require interface length greater than 10, interface-pAE below 10, and interface-pLDDT above 70.
C.1 COMPARISON TO PROTEIN-PROTEIN INTERFACES FROM PROTEIN DATABANK
The section compares six interface metrics between Teddymer and PDB multimer complexes, finding significant distributional overlap that supports Teddymer as training-data augmentation.
- Six metrics compare Teddymer and PDB multimer interfaces: hydrogen bonds, hydrophobicity, shape complementarity, dSASA, and interface residue count.Hydrophobicity is measured separately for the binder interface and binder surface.
- Figure 12 presents the experimental-versus-synthetic interface comparison between Teddymer and PDB multimer.
- Significant overlap between Teddymer and PDB multimer metric distributions supports using Teddymer to augment large-scale binder-generative-model training.The authors report that numerical results and ablations show including Teddymer dramatically boosts performance.
D.1 PDB DATA
The data section describes curated protein, small-molecule, monomer, and synthetic or experimental complex datasets used for training and benchmarking, alongside target-selection and benchmark exclusions.
- PDB data: PDB multimer data were filtered for binder lengths of 50–250 residues and target chains of at least 50 residues before further quality filtering.
- Small-molecule data: PLINDER data were curated for small-molecule binder training using pre-filtered data, sanitized SMILES, and RDKit-processable ligands.
- Monomer data: AFDB monomers used for autoencoder training and latent pretraining required minimum average pLDDT 80 and lengths of 32–256 residues.
- Benchmark targets: The protein-target collection combined AlphaProteo and BindCraft targets into 22 targets, with duplicate PD-L1 targets excluded.
- Benchmark targets: H1, IL17A, and TNF-α were excluded from the main benchmarks because no method achieved success within 20 hours; the final benchmark set contained 19 targets.
- Small-molecule benchmarks: Small-molecule evaluation reused targets from RFDiffusion-AllAtom and BoltzDesign1 and generated 200 structures of length 100 per target.
F EVALUATION METRICS
The evaluation defines protein and ligand success criteria, measures unique successes and novelty, and standardizes sampling-time comparisons across model classes.
- Protein evaluation: Protein designs are successful when ipAE < 7 Å, complex pLDDT > 0.9, and binder scRMSD < 1.5 Å.
- Ligand evaluation: Ligand designs require min ipAE < 2, binder Cα scRMSD < 2 Å, and binder-aligned ligand scRMSD < 5 Å.Ligand scRMSD is computed after aligning the complex to the binder structure.
- Success and novelty metrics: Successes are clustered by Foldseek to report unique successes, while TM-Score against a Foldseek PDB reference set measures structural novelty.Lower TM-Score indicates lower structural similarity and therefore higher novelty.
- Compute evaluation: Sampling time averages 200 samples per target, with a four-hour limit per sample on one NVIDIA A100 GPU.Diffusion models use batch size one; hallucination methods are timed per final optimized design.
- Interface evaluation: Hydrogen bonds are evaluated with HBPlus, and search also uses tmol to calculate interface hydrogen-bond energy.
G ARCHITECTURE, MODEL, TRAINING AND SAMPLING DETAILS
Complexa extends La-Proteína’s partially latent flow-matching model to generate fully atomistic binders conditioned on target structures, with specialized protein and small-molecule representations and training noise that addresses global placement.
- Architecture: Complexa receives a target protein structure and generates a novel fully atomistic binder, while the target serves only as conditioning information.The model processes target and binder jointly, but its generative output is solely the binder.
- Architecture: The denoiser uses pair-biased attention over joint target–binder representations, but predicts the binder’s velocity field rather than generating the target.
- Protein conditioning: Protein representations combine noisy binder coordinates and latent variables with target coordinates, sequence, intra-chain, inter-chain, and hotspot-related pair features.
- Protein conditioning: The feature construction supports multi-chain targets by treating them as one entity and distinguishing chains through a chain-index feature.
- Small-molecule conditioning: Small-molecule conditioning directly featurizes fully atomistic targets because small molecules lack a natural sequence representation.Target features include element types, coordinates, charges, graph positional encodings, and atom names.
- Training and sampling: Training noise and centering are designed to prevent recovering the clean binder’s center of mass trivially and force learning of global binder placement relative to the target.The target is centered at the origin, while translation noise breaks the shortcut exposed by direct binder-noise interpolation.
H INFERENCE-TIME OPTIMIZATION METHOD DETAILS
Complexa scales binder generation at inference time by searching within a learned generative prior using reward-guided sampling and trajectory optimization. These methods improve computational success rates and support challenging targets, while ablations show the value of Teddymer and generative initialization.
- Best-of-N sampling: Best-of-N sampling generates independent designs, evaluates them with a folding model, and retains those meeting predefined success criteria, scaling to 51,200 samples.Generation runs in batches, while folding and scoring use single-sample inference.
- Beam search: Beam search branches each partial trajectory L times, rolls candidates forward to clean samples, rewards them, and retains the top N trajectories.Unlike prior implementations, clean samples used for reward estimation are also added to the success set.
- Feynman–Kac steering: Feynman–Kac steering resamples particles according to reward-derived probabilities, preferentially propagating high-reward trajectories instead of selecting them deterministically.The method samples from a distribution proportional to pϕ(x, s) exp{β R(x, s)}, with β controlling reward bias.
- Ablation findings: Removing Teddymer causes the number of unique successes to drop significantly across most targets, demonstrating its importance for training.The ablation is reported across the target set in Table 7 and Figure 13.
- Hallucination-stage ablation: Logit optimization adds no clear advantage when the base model already supplies a good binder, so compute is redirected toward sequence refinement by mutations.The paper attributes the limited benefit to slow gradients and approximate discrete sequence relaxation.
- Very hard targets: 15 unique TNF-α successes, 7 H1 successes, and 1 IL17A success were identified across 475, 604, and 387 GPU hours, respectively.The corresponding inference-time search scaling curves are reported for these three very hard targets.
I.5 EXTENDED RESULTS: GENERATIVE MODEL BENCHMARK WITH SEQUENCE RE-DESIGN
The benchmark evaluates Complexa and prior models with multiple sequence-redesign strategies across protein and small-molecule targets. Complexa outperforms baselines on most protein tasks, preserves performance with a fixed interface, and produces diverse successful designs.
- Evaluation protocol: 200 samples per model are generated for each of 19 protein targets and evaluated with ProteinMPNN or MPNN-redesign, including fixed-interface and full-sequence settings.Backbone-only models receive eight ProteinMPNN redesign attempts per backbone, while APM and Complexa also use their generated sequences.
- Protein-target results: Complexa clearly outperforms the baselines on most tasks under MPNN-redesign.The comparison is shown in the top plot of Figure 20.
- Interface and sequence redesign: With fixed-interface MPNN-redesign, Complexa performs similarly while APM performs significantly worse, indicating higher-quality generated interfaces for Complexa.For model-generated sequences, the relevant APM baseline collapses for most targets, whereas Complexa retains diverse successes.
- Small-molecule results: Complexa sequences perform well and outperform RFDiffusion-AllAtom across small-molecule targets under the reported LigandMPNN redesign evaluation.RFDiffusion-AllAtom is evaluated only with full sequence redesign because it generates backbones only.
- Result organization: Figure 20 compares unique successes for different models and three sequence-redesign methods across easy and hard targets.Easy targets appear left of the dashed line and hard targets to its right.
I.6 EXTENDED RESULTS: INFERENCE-TIME SEARCH FOR ALL PROTEIN TARGETS
Complexa’s inference-time search methods generally outperform hallucination baselines under matched compute, with simpler sampling favored for easy targets and structured search for harder ones. Reward-guided hydrogen-bond optimization further improves interface quality and often design success.
- Compute-normalized protein-target search: 16 GPU hours for easy targets and 32 GPU hours for hard targets define the maximum inference-time budgets.Budgets include the fixed per-target evaluation setup described for the extended results.
- Easy targets: Best-of-N achieves the highest number of unique successes on 8 of 12 easy targets while consistently outperforming hallucination baselines.Repeated sampling is efficient when the base model already generates many successes.
- Hard targets: Beam Search, Feynman-Kac Steering, and MCTS become advantageous on hard targets where the base model struggles to generate successes.Structured search systematically explores and refines promising candidate regions.
- Limitations and complementarity: On SARS-CoV-2 RBD and HER2-AAV, Complexa yields fewer unique successes than BindCraft, while BoltzDesign is more efficient on BetV1.These exceptions show that hallucination methods retain complementary strengths on some extremely difficult targets.
- Ligand targets: Best-of-N and Beam Search significantly outperform BoltzDesign across the four ligand targets, whereas MCTS performs worst across all targets.The ligand-target results include strong success for SAM despite its difficulty.
- Interface hydrogen-bond optimization: H-Bond-only optimization yields the most interface hydrogen bonds, and combining H-Bond with ipAE consistently improves hydrogen-bond counts over ipAE-only optimization.Across most targets, reward-guided strategies outperform the zero-reward baseline and can also improve unique successes.
I.9 EXTENDED RESULTS: ENZYME DESIGN
Complexa extends binder generation to atomic motif enzyme design and substantially outperforms RFDiffusion2 across the benchmark. Its evaluation uses catalytic-motif, backbone, functional-group, and clash criteria after RF3 refolding, while additional analyses examine scaling and interface realism.
- Benchmark and criteria: The AME benchmark evaluates 41 enzyme-design tasks using catalytic residue recovery, binder-backbone scRMSD ≤2 Å, functional-group scRMSD ≤1.5 Å, and clash avoidance.Clashes are defined as binder atoms within 1.5 Å of ligand atoms.
- Evaluation setup: Complexa’s AME evaluation compares self-generated and LigandMPNN sequences against RFDiffusion2 backbones followed by LigandMPNN sequence generation and RF3 refolding.The comparison also reports best-of-8 LigandMPNN re-designs under the original evaluation protocol.
- Benchmark results: Complexa achieves 41/41 successes with self-generated sequences and 40/41 with a single LigandMPNN re-design, versus RFDiffusion2’s 30/41.With best-of-8 re-designs, Complexa surpasses RFDiffusion2 on 38/41 tasks, including every task with at least four residue islands.
- Long-term scaling: Within 1,100 GPU hours on SpCas9, Complexa generates 13,596 successful binders and 246 unique successes under the default clustering criteria.Unique successes continue increasing with tighter clustering thresholds, while coarser clustering begins to plateau after 600 GPU hours.
- Interface analysis: Reducing ipAE scores correlates with increased binder-target hydrogen-bond interactions, with a Spearman correlation of -0.69.This links lower ipAE values with more interface hydrogen bonding in the reported analysis.
- Interface realism: Generated interfaces generally follow PDB multimer property trends but are slightly less hydrophobic and smaller, with reduced interface shape complementarity.The authors interpret these distributions as indicating realistic target-binder interfaces.
J BASELINE EVALUATIONS
The baseline evaluations document the configurations used for competing methods and provide additional visual evidence for Complexa’s generated binders. These include multi-chain protein targets, small molecules, and fold-class-conditioned designs.
- Baseline configurations: BindCraft, RFDiffusion, Protpardelle-1c, APM, BoltzDesign1, and AlphaDesign are evaluated using publicly released code, checkpoints, or author-provided samples.The evaluation standardizes several sequence re-design and scoring choices across baselines.
- Multi-chain targets: Complexa generates successful de novo binders for TNF-α, IL17A, and H1, including two- and three-chain targets.All shown binders meet the stated in-silico success criteria.
- Small-molecule targets: Complexa generates successful binders for the small molecules SAM, IAI, FAD, and OQO.The visualized designs meet the reported in-silico success criteria.
- Fold-class conditioning: CAT-conditioned Complexa designs follow “Mainly Alpha,” “Mainly Beta,” or “Mixed Alpha Beta” secondary-structure conditioning across five targets.The generated binders exhibit predominantly alpha helices, beta sheets, or both, respectively.
- Model variants: Table 14 distinguishes Protein-Protein, Ligand-Protein, and Protein-Protein CAT Complexa models through their training hyperparameters.Protein-Protein CAT denotes protein-target binders conditioned on CAT labels.