Source-linked AI summary

Representability of algebraic topology for biomolecules in machine learning based scoring and virtual screening

Zixuan Cang, Lin Mu, Guowei Wei

arXiv:1708.08135v1q-bio.QMq-bio.BM

TL;DR

Biomolecular persistent-homology representations must retain chemical, biological, and interaction information while simplifying complex structures. The paper introduces multicomponent, multi-level, and electrostatic persistence with machine learning for molecular characterization, binding-affinity prediction, and virtual screening. Across these tasks, the authors report stronger performance than modern machine-learning methods, while identifying feature-space growth and limited representation of rare elements as important boundaries.

  • Problem

    Persistent homology can lose biological information during geometric simplification, while direct deep CNNs on 3D macromolecules face high cost, insufficient resolution, and inadequate chemical labeling.

  • Method

    The paper constructs multicomponent, multi-level, interactive, and electrostatic persistent-homology representations through tailored distance matrices, then combines them with Wasserstein distances and machine-learning algorithms.

  • Results

    The proposed topological approaches outperform modern machine-learning methods in protein–ligand binding-affinity prediction and ligand–decoy discrimination.

  • Takeaways & Limitations

    Topology-based representations can support both molecular scoring and virtual screening while retaining selected chemical, biological, and interaction information.

  • Takeaways & Limitations

    More element combinations rapidly enlarge the feature space, creating redundant features and potential overfitting.

Abstract

from arXiv · show

This work introduces a number of algebraic topology approaches, such as multicomponent persistent homology, multi-level persistent homology and electrostatic persistence for the representation, characterization, and description of small molecules and biomolecular complexes. Multicomponent persistent homology retains critical chemical and biological information during the topological simplification of biomolecular geometric complexity. Multi-level persistent homology enables a tailored topological description of inter- and/or intra-molecular interactions of interest. Electrostatic persistence incorporates partial charge information into topological invariants. These topological methods are paired with Wasserstein distance to characterize similarities between molecules and are further integrated with a variety of machine learning algorithms, including k-nearest neighbors, ensemble of trees, and deep convolutional neural networks, to manifest their descriptive and predictive powers for chemical and biological problems. Extensive numerical experiments involving more than 4,000 protein-ligand complexes from the PDBBind database and near 100,000 ligands and decoys in the DUD database are performed to test respectively the scoring power and the virtual screening power of the proposed topological approaches. It is demonstrated that the present approaches outperform the modern machine learning based methods in protein-ligand binding affinity predictions and ligand-decoy discrimination.

II Methods and algorithms

Persistent homology represents biomolecular geometry with simplicial complexes and homological invariants, using chains, cycles, boundaries, and quotient groups to characterize connectivity.

  • Simplicial complexes encode discrete biomolecular coordinates through simplices, whose faces and intersections satisfy closure conditions.
  • A k-chain is a formal linear combination of k-simplices over the Z2 field.
  • The boundary operator maps k-simplices or their linear combinations to the corresponding combinations of boundaries in dimension k−1.
  • k-cycles are k-chains with empty boundary, forming the kernel of the boundary operator.
  • The boundary group consists of images of the next-dimensional boundary operator and is contained within the cycle group.

Boundary group

Persistent homology organizes simplicial complexes through nested filtrations and computes Betti numbers from homology groups, which quotient cycles by boundaries.

  • The kth homology group is H_k = Z_k/B_k, and it is used to compute Betti numbers.
  • The kth Betti number is the rank of H_k, equivalently rank Z_k − rank B_k.
  • A filtration is a nested sequence of simplicial subcomplexes from the empty complex to the full complex.
  • Persistent homology tracks homology groups across filtration levels to capture features that persist through scale.
  • Vietoris–Rips complexes include simplices whose pairwise distances do not exceed r, and grow monotonically as r increases.
  • Alpha complexes use Delaunay triangulation and likewise expand monotonically with the scale parameter.

II.B Biological considerations

Biomolecular topology must preserve chemical identity, interaction type, and scale while simplifying complex structures. The proposed representations tailor persistent homology to these biological considerations.

  • Global, intermediate, and local scales may all matter for predictive biomolecular modeling, from conformation to cavities and substructures.
  • Betti-0 barcodes capture covalent-bond lengths, whereas non-covalent interactions require higher-dimensional or tailored representations.
  • Protein–ligand complexes introduce diverse elements beyond the common protein atoms, complicating topological characterization.
  • Multicomponent persistent homology preserves chemical and biological information by selecting element-specific atom combinations.
  • Multi-level persistence and electrostatic persistence target selected interactions, while modified distance matrices extend filtration to interactions not previously covered.

II.D.1 Multi-level persistent homology

Multi-level persistent homology modifies distance matrices to isolate selected intra- or intermolecular interactions and enrich molecular barcode representations beyond covalent-bond information.

  • II.D.1 Multi-level persistent homology: Betti-0 barcodes mainly reflect covalent-bond lengths in small molecules, making hydrogen-bond and van der Waals interactions difficult to capture.
  • II.D.1 Multi-level persistent homology: The modified matrix may violate triangle inequality but still satisfies the construction principle of the Rips complex.
  • II.D.1 Multi-level persistent homology: For the 10-atom 1BCD ligand, ordinary distance matrices produced identical Betti-0-only barcodes for two conformations.
  • II.D.1 Multi-level persistent homology: Using modified matrices enriched the barcode representation with higher-dimensional features that captured small conformational fluctuations.
  • II.D.1 Multi-level persistent homology: An nth-level characterization uses bond-path distance D(i, j) and assigns d∞ beyond the selected bond level.
  • II.D.2 Interactive persistent homology: Interactive persistence retains original distances between two designated atom sets and assigns d∞ to all other pairs.
  • II.D.1 Multi-level persistent homology: Correlation-function filtrations can emphasize geometric scales because interaction strength need not vary linearly with Euclidean distance.

Flexibility and rigidity index based filtration matrix

The filtration framework uses distance- and charge-based correlation functions to encode spatial scale and electrostatic interactions in persistent-homology analyses.

  • Distance-based filtration: Distance-based correlation functions use d_ij and η_ij to separate patterns at different spatial scales.η_ij controls the scale and relates to the radii of two atoms.
  • Electrostatic filtration: Electrostatic filtration rescales interatomic distances using partial charges q_i and q_j and a tunable parameter c.The formulation targets interactions among charged atoms.
  • Electrostatic filtration: Positive c addresses opposite-charge interactions, whereas negative c addresses like-charge interactions.The parameter choice determines which charge relationships are emphasized.
  • Electrostatic filtration: The charge-based correlation regularizes values to the [0, 1] interval for use by machine-learning methods.Weak long-distance or neutral interactions produce correlation values close to 0.5.
  • Alternative filtration: A charge-density construction can also support electrostatic filtration using the charge density as the filtration parameter and a cubical complex.This provides an alternative to the charge-rescaled distance formulation.

II.D.5 Multicomponent persistent homology

Multicomponent persistent homology constructs multiple topological components from combinatorial atom selections, then converts barcode information into machine-learning features.

  • Multicomponent persistent homology: Multicomponent persistent homology systematically generates invariants from different element combinations and constructs 2D or higher-dimensional persistent maps.This differs from element-specific selection, which emphasizes choosing chemically relevant elements for a biological property.
  • Barcode subdivision: Barcode sets can be divided into subsets by slicing persistence diagrams along death, birth, or persistence values.Figure 2 illustrates these three slicing directions.
  • Feature construction: Binned persistent homology counts births, deaths, and persistences within predefined filtration bins.The procedure is applied across Betti-0, Betti-1, and Betti-2 barcodes.
  • Feature construction: Barcode statistics summarize collections using quantities such as averages, standard deviations, extrema, sums, and counts.The feature vectors are built from birth, death, and persistence sets.

Barcode statistics

Barcode statistics provide compact summaries of persistence intervals, while multi-parameter stacking and 2D maps organize these features for convolutional neural networks.

  • Barcode statistics: Statistic feature vectors summarize birth, death, and persistence values across Betti-0, Betti-1, and Betti-2 barcodes.Persistence statistics additionally include the birth and death values of the longest bar.
  • Barcode statistics: Slicing persistence diagrams into subsets enables each subset to receive its own barcode-statistics feature vector.Horizontal slicing groups bars with similar death values and summarizes their birth values.
  • Multi-dimensional persistence: Fixing non-filtration parameters at ordered values and stacking repeated persistence computations provides a tractable multi-dimensional construction.The approach is useful when parameters such as atom selections are discrete and ordered.
  • 2D representation: Eight 2D representations are treated as channels, and protein-ligand analysis adds corresponding complex-minus-protein representations for 16 channels.The channels encode barcode-derived features for element combinations ordered by performance.
  • 2D representation: Each sample is represented by a 120×128×16 tensor from 128 element combinations and 120 bins over [0, 12] Å.These topological images are directly used by deep convolutional neural networks.

II.F Machine learning algorithms

The paper combines KNN, gradient boosting trees, and deep CNNs with barcode-space distances for biomolecular prediction and representation assessment.

  • Algorithms: The machine-learning methods considered are KNN regression, gradient boosting trees, and deep convolutional neural networks.KNN estimates a property from the average or weighted average of its nearest neighbors under a dataset metric.
  • Barcode-space metrics: Wasserstein metrics measure similarity between barcode sets generated from different biomolecules.The Wasserstein metric is an Lp generalization of the bottleneck distance, with p = 2 chosen here.
  • Scope: An exhaustive comparison with other metric-space distances, including Hausdorff and Gromov-Hausdorff distances, is beyond the work’s scope.The paper notes these distances as possible alternatives for biomolecular analysis.
  • Representation assessment: Barcode-space metrics assess persistent-homology representation power without potential overfitting from manually generated feature vectors.The authors report that metric-induced similarity is significantly correlated with molecule functions.
  • Scope: Wasserstein metric measures could be used directly in kernel methods such as nonlinear support vector machines, but this option is not explored.The stated application includes classification and regression tasks.

II.F.2 Gradient boosting trees

The study uses gradient boosting trees and deep convolutional neural networks within topology-based scoring and virtual-screening workflows. The CNN architecture and training setup are specified, while docking poses can be rescored to rerank candidates.

  • Gradient boosting trees: Gradient boosting trees ensemble decision trees to learn complex feature–target maps while reducing overfitting through shrinkage.The implementation uses scikit-learn's GradientBoostingRegressor.
  • Deep convolutional neural networks: The deep CNN begins with convolution layers followed by dense layers and uses manually assigned parameters without optimization.It uses Adam with learning rate 0.0001, mean squared error, batch size 32, and 500 epochs.
  • Deep convolutional neural networks: Figure 4 encodes convolution filters, filter sizes, activations, dense-layer neurons, pooling sizes, dropout rates, and repeated layers.Structured layers appear in boxes, unstructured layers in rectangles, and repeated layers carry a ×2 marker.
  • Applications: The topology-based workflow supports scoring and virtual screening, including rescoring docking-generated poses to rerank candidates.The paper applies the method as a rescoring machine based on docking poses.

III.A Ligand based protein-ligand binding affinity prediction

Ligand-based experiments evaluate whether persistent-homology representations predict binding affinity across seven protein clusters. They compare element-selected topological features, multi-level distances, Wasserstein-based nearest-neighbor regression, and feature representations.

  • Dataset and setup: 1322 protein-ligand complexes across 7 protein clusters form the S1322 benchmark, using ligand structures alone for topological analysis.The dataset is a subset of the PDBBind v2015 refined set.
  • Topological representations: The experiments compute Rips and alpha persistent homology through Betti-2 for atom collections selected by element type and distance construction.Features combine barcode statistics or counts in bins, with multi-level distance matrices included for Rips computations.
  • Evaluation: Binding-affinity prediction uses repeated 10-fold or 5-fold cross-validation, reporting median Pearson correlation coefficients and RMSE in kcal/mol.The repeated validation is performed 20 times for each feature-vector subset.
  • Wasserstein-based regression: Wasserstein distances between barcode sets support leave-one-out k-nearest-neighbor regression with k = 3 within each protein cluster.Figure 7 compares the best- and worst-performing clusters by the proximity of neighbor connections to the semicircle.
  • Results: Topological features based on barcode statistics typically outperform features based on counts in bins.Multi-level persistence significantly outperforms Euclidean-distance persistence in protein clusters 3 and 6.

Performance of multicomponent persistent homology

The paper evaluates multicomponent persistent-homology representations for ligand-based and protein-ligand-complex binding-affinity prediction using PDBBind benchmarks. Ligand-only topological descriptors show greater predictive power than earlier physical descriptors, while feature redundancy is generally tolerated.

  • Comparative performance: Ligand-only topological descriptors have more predictive power than physical descriptors built from protein-ligand complexes in earlier work.The comparison uses FFT-BP results based on multiple additive regression trees and physical descriptors.
  • Evaluation scope: The ligand-based experiments cover 7 protein clusters and 1322 complexes, while additional experiments evaluate protein-ligand-complex representations on PDBBind datasets.The study summarizes ligand-based experiments in Table 1 and complex-based experiments in Table 4.
  • Feature construction: Rare elements are retained because omitting them could reduce the model's ability to handle new data where those elements matter.The trade-off is increased feature count and redundancy, exemplified by rare bromine occurrences.
  • Feature construction: Feature accumulation tests show that the model is robust against adding more element combinations in most cases.Performance is measured with Pearson correlation coefficient as element-combination features are added incrementally.
  • Datasets: PDBBind supplies curated protein-ligand structures and binding-affinity data and serves as a benchmark resource for computational binding analysis.The study uses four PDBBind core sets as test sets, with corresponding refined sets excluding test complexes for training.
  • Evaluation: Table 5 reports Pearson correlation coefficients with RMSE for feature-group comparisons across four PDBBind core sets.Ensemble-of-trees results use medians over 50 repeated runs, while deep-learning methods use consensus models from independently generated models.

Robustness of GBT algorithm against redundant element combination features and potential overfitting

The study tests whether expanding element-combination features improves characterization without making GBT models overly sensitive to redundant information. Across PDBBind analyses, performance remains stable as combinations increase, while features and training samples for rare elements remain important.

  • Feature-combination design: 128 element combinations are generated by pairing protein and ligand element-group choices, then added incrementally in importance-score order.The procedure evaluates progressively larger feature vectors, beginning with the most important combinations.
  • Robustness to redundant features: Top element combinations readily produce good PDBBind models, while performance fluctuates within a small range as additional combinations are included.The v2007 trend is nearly monotonic; the other three data sets show unsteady but limited variation.
  • Robustness to redundant features: GBT algorithms are not very sensitive to redundant features and remain robust against potential overfitting when model parameters are fixed.This conclusion follows from the observed stability across increasing numbers of element combinations.
  • Rare-element robustness: For most rare elements, using all features with the original training set performs best, whereas excluding matching training samples performs worst.Figure 10 averages RMSE across the PDBBind v2007, v2013, v2015, and v2016 core sets.
  • Interactive Betti-0 combinations: Interactive Betti-0 tests compare 36 basic element combinations with 160 selected combinations to assess whether larger combinations add useful information.The comparison specifically addresses redundancy among higher-order interactive Betti-0 barcodes.
  • Electrostatic information: Explicitly embedding atomic charges generally outperforms Euclidean-distance features alone, indicating that electrostatics matter for protein-ligand binding prediction.Charges are incorporated into interactive Betti-0 barcodes through the electrostatic persistence construction.

2D persistence for deep convolutional neural networks

The work converts multicomponent persistent-homology information into a 2D, 16-channel representation for deep convolutional neural networks. CNNs significantly outperform GBTs on three larger PDBBind data sets, while the smallest set is an exception.

  • 2D representation: 128 element combinations become an additional representation dimension, producing 16 channels for deep-learning input.The channels encode topological information from multicomponent persistence for CNN processing.
  • Performance across data sets: For PDBBind v2013, v2015, and v2016, the 2D CNN representation performs significantly better than the alternative model.PDBBind v2007 is the only reported exception.
  • Performance across data sets: The inferior CNN performance on PDBBind v2007 may result from its smaller training set of 1,102 complexes versus more than 2,700 for the other sets.The reported comparison links the data-size difference to the observed performance pattern as a possible explanation.
  • Virtual-screening setting: Virtual screening uses DUD structures for 40 protein targets, with docking-generated protein-ligand and protein-decoy complexes.The evaluation distinguishes active ligands from decoys using structure-based complex data.
  • Virtual-screening metrics: AUC and enrichment factor evaluate how effectively each method ranks active ligands above decoys.AUC ranges from random selection at 0.5 to perfect prediction at 1.

Topology based machine learning models

The paper develops topology-based machine-learning models to address the limited representability of persistent homology for small molecules and macromolecular interactions. Across binding-affinity and virtual-screening applications, the proposed descriptors support competitive or improved predictive performance, with complex-based features sensitive to docking quality.

  • Proposed topology-based models: TopVS-ML combines manually constructed topological features with gradient boosting, random forest, and extra-trees voters.The final virtual-screening model uses descriptors for both protein-compound interactions and compounds alone.
  • Virtual screening: An AUC of 0.83 is obtained by combining compound-only and protein-interaction descriptors, exceeding compound-only AUC 0.81 and interaction-only AUC 0.77.The reported result treats the two descriptor groups as complementary.
  • Docking-quality dependence: For low-quality docking with Autodock Vina AUC < 0.5, ligand-based features achieve AUC 0.81 while complex-based features achieve AUC 0.74.For high-quality docking with Autodock Vina AUC > 0.8, the corresponding values are 0.81 and 0.86.
  • Virtual screening: The small-molecule topology model reaches AUC 0.81 and is reported as competitive with other leading methods.The paper also reports competitiveness in binding-affinity prediction using experimentally solved complexes.
  • Motivation and scope: The work addresses an earlier gap because the representability and predictive power of element-specific persistent homology for small molecules and macromolecular interactions were unknown.The proposed methods are evaluated across more than 4,000 PDBBind complexes and nearly 100,000 DUD compounds.
  • Proposed topology-based models: The paper introduces multicomponent, multi-level, and electrostatic persistence to represent chemical composition, selected interactions, and partial charges.These methods construct distance matrices for filtration and are evaluated on PDBBind and DUD.
Loading 1708.08135v1…