Source-linked AI summary
Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN
Eliuvish Han Cui
TL;DR
Scarce PET confirmation must support prespecified claims, but generic uncertainty sampling may prioritize subjects with little target influence. The paper uses A4/LEARN as a retrospective design laboratory to compare target-aligned and simpler validation designs. For the primary APOE4 contrast, balanced validation nearly matches target-specific scoring, while other targets show larger gains from target alignment.
Problem
PET confirmation remains scarce even as cheaper first-phase information broadens Alzheimer disease screening, creating a need to choose which subjects support the reported claim.
Method
The paper emulates scarce-PET validation in the A4/LEARN archive, comparing APOE4-balanced and target-specific designs with generic uncertainty sampling.
Results
At PET budget 200, the confidence-interval width ratio is 0.923 for APOE4-balanced validation, 0.914 for target-specific scoring, and 0.980 for generic uncertainty sampling.
Takeaways & Limitations
Spend scarce protocol measurements according to the claim being validated: simple APOE4 balancing is nearly enough for the primary contrast, but target-specific scoring helps more for uneven or cutoff-indexed targets.
Takeaways & Limitations
The analysis is retrospective, using the complete PET archive as a design laboratory rather than demonstrating prospective deployment.
Abstract
from arXiv · showhide
Anti-amyloid therapies and blood-based biomarkers are changing Alzheimer disease workups into a two-stage measurement workflow: screen broadly with cheaper information, then spend scarce confirmatory amyloid measurements where they support the decision that will be reported. Amyloid positron-emission tomography (PET) remains one such protocol measurement for amyloid burden, but PET slots, trial budgets, and payer-facing evidence packages are finite. This paper asks a deliberately operational question: when is simple transparent PET validation enough, and when is a fitted residual-uncertainty score worth the added complexity? For a weighted protocol target, the first-order value of validating subject i is the product of target influence and residual protocol uncertainty. Generic uncertainty sampling uses only the second factor and can spend PET measurements on subjects that are hard to predict but weak for the scientific, clinical, or commercial claim. We apply this rule to the A4/LEARN PET archive, treating observed PET as a design laboratory for scarce-confirmation studies. For the primary APOE4 carrier versus non-carrier contrast in Centiloid 24-or-higher PET positivity, simple APOE4-balanced validation recovers nearly all of the target-specific gain: at PET budget 200, the confidence-interval width ratio relative to random validation is 0.923 for APOE4 balancing and 0.914 for target-specific scoring, while generic uncertainty sampling is 0.980. Other targets behave differently: target-specific scoring gives larger gains for an age-slope analysis and for cutoff-indexed PET positivity. The practical message is simple: spend scarce protocol measurements according to the claim being validated, not only according to prediction uncertainty.
1 Why scarce confirmatory measurement is the problem
Cheaper first-phase information broadens screening, but scarce PET confirmation still must be allocated around the claim being validated. The paper contrasts generic uncertainty sampling with transparent or target-aligned designs for APOE4 and other PET-defined targets.
- Why scarce confirmatory measurement is the problem: Broader screening does not remove the PET bottleneck when PET is the protocol measurement anchoring amyloid burden.The validation budget may be smaller than the screened population.
- Why scarce confirmatory measurement is the problem: Different APOE4, age, cutoff, and treatment-eligibility targets need not be served by the same validation subset.The relevant subset depends on the question the PET measurements are meant to answer.
- Why scarce confirmatory measurement is the problem: Generic uncertainty sampling can prioritize subjects who are difficult to predict even when they have weak influence on a prespecified claim.Its appeal is strongest when the objective is improving the predictor, not necessarily supporting a specific evidence package.
- Why scarce confirmatory measurement is the problem: For the APOE4 carrier versus non-carrier PET-positivity contrast, transparent APOE4-balanced validation performs almost as well as fitted target-specific validation.This is the paper’s main empirical lesson for the primary target.
- Why scarce confirmatory measurement is the problem: The same near-equivalence does not automatically transfer to age slopes, cutoff-indexed PET endpoints, or cohort-structured diagnostic contrasts.These targets motivate target-specific scoring rather than assuming one design works universally.
- Why scarce confirmatory measurement is the problem: The contribution is a decision rule that uses surrogate information without allowing the surrogate to redefine the target.The operational sequence is to define the claim, anchor it to the protocol measurement, record probabilities before outcome use, and allocate scarce measurements accordingly.
2 The A4/LEARN validation decision
A4/LEARN is used as a retrospective design laboratory for deciding how scarce PET confirmations should support finite-archive targets. The comparison includes random, generic-uncertainty, APOE4-balanced, and target-specific validation rules under documented data support.
- The A4/LEARN validation decision: The analysis uses the July 2024 A4/LEARN controlled-access release and emulates PET validation by hiding outcomes before spending a validation budget.Observed PET is then used to evaluate finite-archive inference for prespecified targets.
- The A4/LEARN validation decision: The archive contains 6,945 rows, including 4,492 with nonmissing Centiloid PET and 4,460 with both PET and APOE4.
- The A4/LEARN validation decision: The primary target is the finite-archive difference between APOE4 carriers and non-carriers in CL24 positivity, not a universal disease boundary.Population language is shorthand for this empirical contrast rather than an external-validity claim.
- The A4/LEARN validation decision: The main claims concern validation design under documented feature support, not clinical deployment of a particular PET-risk model.Feature availability varies across cognitive, plasma, pTau217, and MRI summaries.
- The A4/LEARN validation decision: The four comparison designs are random validation, generic uncertainty sampling, APOE4-balanced validation, and target-specific allocation by target influence times residual uncertainty.
- The A4/LEARN validation decision: Budget 200 is a readable mid-range scenario, and all reported ratios are relative to random validation at the same PET budget.Oracle lower-bound rows are infeasible diagnostics, not deployable designs.
3 Target-aware rule and recorded estimator
The paper defines validation value using both target influence and residual protocol uncertainty, while recording inclusion probabilities before PET outcomes enter the estimator. This distinguishes target-aligned sampling from generic uncertainty sampling.
- Target-aware rule and recorded estimator: Target influence separates APOE4 carriers from non-carriers, reflects age leverage for an age slope, and changes with the PET threshold for a cutoff curve.
- Target-aware rule and recorded estimator: Residual protocol uncertainty is the uncertainty remaining after first-phase information is used to predict the PET-defined quantity.The first-phase prediction and residual uncertainty are denoted by µ_i(t) and σ_i(t).
- Target-aware rule and recorded estimator: The estimator records each positive validation probability π_i before PET outcomes enter the final correction, with R_i indicating validation.This preserves estimator centering through recorded probabilities and positivity rather than requiring a correct surrogate model.
- Target-aware rule and recorded estimator: For a scalar target, the first-order value of validating subject i is proportional to target influence times residual protocol uncertainty.The target weight encodes what the study is trying to estimate.
- Target-aware rule and recorded estimator: Generic uncertainty sampling uses σ_i(t), whereas target-aligned validation uses |a_i(t)|σ_i(t).The difference matters when prediction uncertainty and target influence identify different subjects.
- Target-aware rule and recorded estimator: In prospective deployment, residual uncertainty must come from information available before the final PET validation decision; full-PET residuals are retrospective evaluation tools here.
4 Primary APOE4 result: simple balancing is nearly enough
For the APOE4 carrier-versus-non-carrier target, APOE4-balanced validation nearly matches fitted target-specific scoring, while generic uncertainty sampling adds little. The result supports transparent balancing as the practical default comparator for this estimand.
- The cross-fitted surrogate AUC is about 0.779, but AUC is not the design conclusion.The design conclusion concerns validation precision rather than predictive discrimination alone.
- APOE4 balancing recovers most of the target-specific gain for the primary two-group contrast.APOE4 status defines the contrast, so ensuring measurement in both strata captures the key influence structure before fitted scoring.
- 0.914 is the CI-width ratio for target-specific scoring at PET budget 200, versus 0.923 for APOE4-balanced validation and 0.980 for generic uncertainty sampling.All ratios are relative to random validation.
- Balanced validation is easier to explain, audit, and implement than residual-uncertainty scoring.The paper recommends it as the serious default comparator for this target, not as a universal replacement for fitted scores.
- A stress test excluding APOE4 from the PET-risk surrogate produced AUC 0.724, surrogate-only contrast 0.069, PET contrast 0.334, and bias -0.265.APOE4 remained the prespecified target-defining variable in this stricter score.
5 Scope checks: when targeting matters
Target-specific scoring matters more when influence is heterogeneous or the target varies across an index such as age or PET cutoff. The A4-versus-LEARN analysis further shows that cohort structure can distort surrogate-based contrasts.
- Age-slope target: 0.806 is the CI-width ratio for target-specific validation of the age-slope target at budget 200, versus 1.000 for generic uncertainty sampling and 1.001 for age-quartile balancing.Coarse balancing does not equal allocation by leverage and residual protocol uncertainty.
- Cutoff-indexed positivity: 0.859 is the CI-width ratio for cutoff-range design across thresholds 20 to 30, compared with about 0.905 for generic uncertainty validation.At c = 24, the PET contrast is 0.334 and the threshold-specific surrogate contrast is 0.347.
- A4-versus-LEARN diagnostic layer: The A4-versus-LEARN comparison is diagnostic rather than a causal or biological cohort contrast.Its value is to show that surrogate error can align with cohort structure; it should not dominate the primary biomedical story.
- A4-versus-LEARN diagnostic layer: Adding cohort information raises AUC to about 0.908 and reduces contrast-aligned bias from -0.493 to -0.039.Without cohort information, the surrogate-only A4-versus-LEARN contrast is 0.380 versus the PET protocol contrast of 0.873.
6 Deployable pilot evidence
A deployable design uses a random pilot to learn target-specific allocation for a second PET wave. Pilot emulation improves precision without requiring unobserved PET outcomes when choosing that second wave.
- The two-wave design is more practice-relevant because it does not require unobserved PET outcomes to choose the second wave.
- At total budget 400 with a random pilot of 100, the corresponding ratios are 0.917 and 0.918, with coverage 0.950.These results are less dramatic than full-archive oracle diagnostics.
- A richer learner may reduce prediction error, but safe-feature scores support the same qualitative APOE4 conclusion.
- pTau217 sensitivity is kept outside the main design because its archive timing is not uniformly pre-PET.
7 Discussion
The discussion recommends aligning scarce PET validation with the estimand: transparent APOE4 balancing is nearly as effective for the primary contrast, while target-specific scoring is preferable for more uneven or cutoff-indexed targets. The retrospective design and heterogeneous first-phase support constrain prospective interpretation and deployment.
- Practical planning: Target-specific validation aligns allocation with target influence and residual uncertainty rather than prediction uncertainty alone.Prospective use requires learning residual uncertainty from historical data or a random pilot and recording validation probabilities before outcome observation.
- Practical planning: APOE4-balanced validation is transparent and nearly matches target-specific design for the primary APOE4 PET contrast.For targets with uneven influence, cutoff-indexed structure, or within-stratum residual heterogeneity, the discussion recommends target-aligned residual-uncertainty scoring.
- Boundaries: The analysis is retrospective: the complete PET archive evaluates scarce-validation designs but does not claim future PET outcomes would be observed under those designs.The discussion also places crossed stratified grids and cutoff-aware balanced validation outside the current performance package.
- Measurement discipline: 92.2% agreement across 4,492 paired rows supports PET visual read as a plausibility audit, not a replacement endpoint.Cohen’s kappa was 0.821, while PET remains the protocol-defined measurement.
- Measurement discipline: Plasma, MRI, pTau217, and cognitive features should guide targeting only when available in the intended first-phase cohort because their timing and support are heterogeneous.The paper keeps auxiliary information subordinate to the protocol quantity rather than replacing it.
- Scope and implementation: The current empirical claim concerns scarce confirmatory PET measurements for scalar and cutoff-indexed amyloid targets, while time-indexed extensions remain outside the needed analysis.The discussion emphasizes documenting the target, first-phase support, validation probabilities, and final PET-defined estimator.
Data availability
Individual A4/LEARN participant-level data are controlled-access study data available to qualified researchers through the A4 and LEARN Study Data Portal. Reproducibility materials provide manuscript and analysis resources without redistributing controlled-access source files.
- Access: Qualified researchers can access individual A4/LEARN participant-level data through A4StudyData.org subject to registration, approval, and data-use requirements.Controlled-access source files are not redistributed.
- Reproducibility: The reproducibility materials include manuscript sources, analysis and simulation code, cleared aggregate tables and figures, requirements information, and synthetic data.
Use of AI/NLP tools
OpenAI Codex and ChatGPT assisted with drafting, code organization, formatting checks, and submission preparation. The author reviewed the work and retains responsibility for its content, code, analyses, and conclusions.
- Declared use: OpenAI Codex and ChatGPT were used as editorial and programming assistants for drafting, code organization, formatting checks, and submission preparation.
- Responsibility: The author reviewed the manuscript and takes responsibility for all content, code, analyses, and conclusions.