Source-linked AI summary

Predicting Quantifiability from Primary Screens to Prioritize Dose-Response Profiling

Sean Lim

arXiv:2608.26538v1cs.LGq-bio.BM

TL;DR

Dose-response profiling is costly, and biological activity does not guarantee a usable potency estimate. This paper models quantifiability from primary-screen information and finds that screening-response features strongly predict reportable potency, supporting quantifiability-aware allocation of profiling capacity.

  • Problem

    Biological activity does not always yield a reliable, reportable potency estimate, creating a distinct quantifiability problem for costly follow-up profiling.

  • Method

    The study models quantifiability for 32,971 historically promoted compound–assay pairs using screening responses and complementary molecular, historical, and assay information.

  • Results

    Quantifiability is strongly predictable from screening responses, generalizes across unseen scaffolds and assay-mechanism families, and varies with response amplitude and assay context.

  • Takeaways & Limitations

    Quantifiability-aware ranking can concentrate dose-response profiling on pairs more likely to yield potency values, improving allocation of profiling capacity within the evaluated workflow.

  • Takeaways & Limitations

    The analysis is retrospective and conditional on the historical promotion policy, and transfer to another laboratory requires local validation and possible recalibration.

Abstract

from arXiv · show

High-throughput drug screening relies on low-cost primary assays to prioritize compounds for more expensive dose-response profiling, where potency is ultimately quantified. Current screening strategies largely focus on identifying compounds that will confirm biological activity on follow-up, implicitly assuming that confirmed activity will also yield a usable potency estimate. However, confirmed biological activity in screening does not necessarily translate into a quantifiable potency, because active compounds can still fail to produce a reportable dose-response estimate. We therefore present a framework for modeling quantifiability, whether follow-up testing will yield a usable potency estimate, as a distinct triage objective from biological activity. Quantifiability was strongly predictable from the preceding low-cost screen, with most predictive information arising from the observed screening features rather than molecular structure. Response-based predictors remained robust on previously unseen chemical scaffolds and generalized across held-out assay-mechanism families, while the probability of successful quantification varied strongly with response amplitude and assay context. These findings establish experimental measurability, distinct from biological activity, as a predictable property of screening outcomes and show that quantifiability-aware triage can improve the allocation of costly dose-response profiling capacity.

1 Introduction

Dose-response profiling is costly, yet biological activity alone does not ensure a reportable potency estimate. The paper therefore proposes predicting quantifiability from primary-screen information to prioritize profiling.

  • Potency estimates from concentration–response experiments support structure–activity relationships, selectivity profiles, and progression decisions.
  • Primary screens are inexpensive and broad, whereas multi-concentration profiling is resource-intensive, making promotion a critical allocation decision.
  • Conventional triage emphasizes whether screening signals reflect genuine activity, noise, or assay interference, primarily targeting false-positive activity.
  • Active compounds can still produce weak or incomplete responses that fail to support a reliable Hill fit and reportable pXC50.
  • The study formulates quantifiability as a distinct triage target and evaluates it across 32,971 promoted compound–assay pairs spanning five assay-mechanism families.
  • Quantifiability predictions rely mainly on screening curves: 16 three-point features outperform traditional heuristics, while Morgan fingerprints lose apparent signal under scaffold-disjoint holdout.
  • Failure rates range from 18.8% to 81.6% across assay formats, showing that the response amplitude needed for quantification depends on assay context.

2 Methods

The study predicts reportable potency from promoted compound–assay pairs using screening responses, molecular fingerprints, historical features, and assay metadata. It evaluates discrimination, scaffold and mechanism generalization, chronological deployment, and profiling-budget metrics.

  • Study design and outcome definition: The dataset contains 32,971 pairs promoted from three-point screens to 11-point dose–response experiments, with quantifiability defined as producing a reportable pXC50.
  • Predictive representation: The model uses 16 three-point response features, 1,024-bit Morgan fingerprints, historical compound and assay features, and assay metadata.
  • Predictive models: A class-balanced random forest predicts the binary quantifiability outcome, with logistic regression and isolated feature blocks providing comparators.
  • Generalization analyses: Nested feature-group ablations, five-fold scaffold-disjoint splitting, and leave-one-mechanism-family-out evaluation test information-source contributions and generalization.
  • Chronological evaluation: Leakage-safe rolling evaluation scores each release using only earlier releases and recomputes historical features strictly before the test release.
  • Operational ranking: The operational score factorizes quantifiability through activity and conditional quantifiability, while predictions are evaluated as within-release rankings rather than fixed probability thresholds.
  • Profiling-policy metrics: Recall@50% measures recovered quantifiable pairs among the top half of a ranking, while yield measures precision in that same selected prefix.
  • External validation: External validation tests the amplitude–quantifiability association on matched EPA ToxCast/Tox21 compound–endpoint pairs but does not validate the EvE three-point model or its budget reductions.

3 Results

Quantifiability was predictable from three-point screening data, with the strongest information coming from observed response features rather than molecular structure. Weak or incomplete responses and assay context shaped whether active pairs yielded reportable potencies, and quantifiability-aware ranking improved profiling efficiency in chronological evaluation.

  • 3.1 Quantifiability and potency can be predicted from the three-point screen: 32,971 promoted pairs produced 14,672 reportable potencies, while 6,750 active profiles remained unquantifiable.The 6,750 active-unquantifiable profiles represented 20.5% of all promoted pairs and 31.5% of pairs confirmed active during profiling.
  • 3.1 Quantifiability and potency can be predicted from the three-point screen: Mean three-point activity achieved an AUROC of 0.863 ± 0.002, establishing a strong screening-activity baseline for quantifiability ranking.Maximum three-point activity achieved 0.848 ± 0.003 under the same five-fold random cross-validation.
  • 3.1 Quantifiability and potency can be predicted from the three-point screen: The full random forest reached an AUROC of 0.939 ± 0.002 and recovered 89.4% of quantifiable pairs at a 50% profiling budget.Performance varied less across model classes than across feature sets, indicating that screening information drove most of the predictive performance.
  • 3.2 Predictive information is concentrated in the screening response rather than molecular structure: The three-point feature block achieved an AUROC of 0.850 ± 0.005 versus 0.688 ± 0.010 for Morgan fingerprints alone.Adding fingerprints produced the largest subsequent nested gain, while historical profiles and assay metadata added smaller gains.
  • 3.2 Predictive information is concentrated in the screening response rather than molecular structure: Under scaffold-disjoint splitting, the three-point-only AUROC changed from 0.886 ± 0.002 to 0.883 ± 0.021, while fingerprint-only AUROC fell from 0.761 ± 0.004 to 0.599 ± 0.042.The full model changed from 0.939 ± 0.002 under random splitting to 0.919 ± 0.017 under scaffold-disjoint splitting.
  • 3.3 Active but unquantifiable profiles are characterized by weak response amplitude: Mean three-point maximum activity was 68.0% for quantifiable pairs, 42.6% for active-unquantifiable pairs, and 19.0% for pairs inactive during profiling.Hill-fitting failure ranged from 18.8% in G-protein activation assays to 81.6% in heterodimer assays, showing assay-context dependence.
  • 3.4 Chronological evaluation quantifies the profiling-budget trade-off: Rolling-evaluation AUROC ranged from 0.814 to 0.977, and the model recovered more quantifiable pairs than screening-activity ranking at a 50% budget on all eight releases.Across chronological test releases, profiling 55.4% of promoted pairs recovered 90% of potencies obtained by exhaustive profiling, with a smaller budget than the screening heuristic on seven releases.

4 Discussion

The study reframes dose–response triage around quantifiability—the likelihood that profiling returns a usable potency estimate—rather than biological activity alone. Screening responses predict this outcome, but savings and thresholds depend on assay context and incoming-release composition.

  • More than one-third of unsuccessful profiles were active at 11 concentrations but still failed to yield a reportable potency.
  • Three-point screening features substantially outperformed molecular fingerprints alone and retained performance on unseen chemical scaffolds and held-out assay-mechanism families.
  • Active but unquantifiable profiles had lower amplitudes than quantifiable profiles, while failure rates ranged from 18.8% to 81.6% across assay-mechanism families.
  • At a 90% retention target, retrospective ranking retained 90% of exhaustive-profile potencies while profiling 55.4% of promoted pairs.
  • The achievable profiling budget varied with the prevalence and composition of quantifiable pairs in each release, so the model defines a tunable budget–retention trade-off rather than a universal savings rate.
  • The analysis is retrospective and conditional on historically promoted pairs; transfer to another laboratory requires local validation and possibly recalibration, with prospective evaluation still needed.

Data Availability.

The study uses publicly available datasets from EvE Bio and EPA ToxCast/Tox21 under the stated access conditions.

  • Primary analyses used publicly available EvE Bio Data Releases #1–#11 containing three-point screening and 11-point concentration–response data.
  • External validation used publicly available U.S. EPA ToxCast/Tox21 high-throughput screening and concentration–response data.
  • The EvE Bio data are available through the EvE Bio Data portal under a Creative Commons CC BY-NC-SA 4.0 license.
  • The author declares no competing interests.
Loading 2608.26538v1…