Source-linked AI summary

Anchored Scenario Coverage for Failure-Aware First-Hit Batch Inverse Design

Chuhan Yang, Chenxi Wang, Linhan Wu, Yuyang Liu

arXiv:2608.27873v1math.OCcs.LG

TL;DR

Failure-prone closed-loop inverse design needs early discovery of at least one valid target design, but independently selecting high marginal-score candidates can waste batch capacity through predictive redundancy. ARC-SC retains strong marginal candidates as risk-admissible anchors and diversifies the remaining positions through predictive scenario coverage; frozen-oracle simulations support earlier first-hit discovery, especially under structured failure and redundancy.

  • Problem

    Independently selecting high marginal valid-hit candidates can produce redundant batches under predictive uncertainty, weakening first-hit discovery when shared predictive assumptions fail.

  • Method

    ARC-SC anchors candidates using the marginal valid-hit criterion, then fills the remaining batch positions by maximizing complementary predictive target-scenario coverage under risk-support control.

  • Results

    Across frozen-oracle superconductivity and JARVIS materials benchmarks, ARC-SC produces more earlier first-hit outcomes than later outcomes relative to ordinary batch POF, with strongest gains under feature-interaction and low-validity failures.

  • Takeaways & Limitations

    ARC-SC is supported as a POF-preserving diversification strategy for settings where structured experimental failure and predictive redundancy matter.

Abstract

from arXiv · show

Early discovery of at least one valid design satisfying a target requirement is a central objective in failure-prone closed-loop inverse design. A natural batch baseline ranks candidates by a product-form marginal valid-hit score, but selecting the highest-ranked candidates independently can produce redundant recommendations under predictive uncertainty and waste the experiment budget. We introduce ARC-SC(Anchored Risk-Constrained Scenario Coverage), a batch acquisition method that preserves strong marginal candidates as anchors and allocates the remaining batch positions by maximizing complementary coverage over predictive target scenarios under a risk-support constraint. In frozen-oracle closed-loop simulations on superconductivity and JARVIS materials-property benchmarks, ARC-SC yields a statistically supported improvement in first-hit discovery and remains competitive with directionally favorable first-hit performance on more challenging design space. These results establish ARC-SC as a POF-anchored, scenario-aware batch strategy for improving early valid-target discovery under structured experimental failure.

1 INTRODUCTION

Failure-aware closed-loop inverse design prioritizes finding an early valid target hit rather than optimizing only the final objective. ARC-SC addresses batch redundancy by anchoring strong marginal candidates and diversifying remaining selections across predictive scenarios under risk control.

  • Motivation: First-hit discovery seeks at least one candidate satisfying validity and target requirements with as few costly evaluations as possible.This objective matters especially when evaluations are expensive, time-consuming, destructive, or performed in parallel batches.
  • Baseline: The POF baseline ranks candidates independently by the product of target-satisfaction and valid-outcome probabilities.Under calibrated marginal probabilities and conditional independence, this ranking maximizes the current-round probability of at least one successful design.
  • Batch gap: Shared predictive explanations can make high-scoring candidates redundant, leaving the batch vulnerable when the common explanation is wrong.This redundancy is especially consequential because first-hit value depends mainly on whether at least one batch member succeeds.
  • ARC-SC: ARC-SC preserves strong marginal valid-hit candidates as anchors, then fills remaining positions through predictive scenario coverage under risk-support control.Its construction combines marginal valid-hit anchoring, scenario coverage, and batch-level risk admissibility.
  • Theory: The scenario-coverage objective is normalized, monotone, and submodular, yielding the classical fixed-domain greedy-completion guarantee.The guarantee applies with fixed anchors and a fixed admissible completion pool; the full implementation instead uses a dynamic risk constraint.
  • Evaluation: Across two materials benchmarks and controlled failure settings, ARC-SC produces more earlier first-hit outcomes than later outcomes relative to ordinary batch POF.Its clearest gains occur under feature-interaction and low-validity failures, while many hybrid and off-manifold configurations tie.

2 METHODOLOGY

ARC-SC formulates failure-aware closed-loop inverse design around early discovery of a valid candidate meeting a target threshold. It combines marginal valid-hit exploitation with risk-constrained, scenario-aware batch completion to reduce redundancy under predictive uncertainty.

  • Candidate modeling: ARC-SC updates target and validity models from observed valid and failed evaluations, then scores a finite candidate pool.Target predictions use valid observations, while validity predictions use successful and failed evaluations; empirical support decreases for extrapolative candidates.
  • Problem formulation: The optimizer seeks the earliest batch containing a candidate that is both valid and above the user-defined target threshold.First-hit discovery is evaluated by the first-hit round, with no observed hit assigned an infinite hitting time.
  • Scenario coverage: ARC-SC fills remaining batch positions with candidates that provide complementary coverage across multiple predictive target scenarios.Scenario responses combine marginal valid-hit, validity, empirical-support, and soft target-response information; scenario construction uses marginal predictive uncertainty on the candidate pool.
  • Risk-constrained completion: A dynamic risk-support constraint limits average batch risk during completion, while a penalized fallback completes the batch when no feasible extension exists.The support score is high near observed designs and declines for candidates farther from the observed design distribution.
  • POF-anchored selection: The method first selects risk-admissible POF anchors using marginal valid-hit scores, preserving strong candidates for immediate success.Because anchoring obeys a dynamic risk constraint, the anchors need not equal the first K independently ranked POF candidates.

3 THEORETICAL ANALYSIS

ARC-SC's scenario-coverage objective has a formal diminishing-returns structure, yielding a greedy guarantee for fixed anchors and a fixed admissible completion pool. The full dynamically risk-filtered procedure remains outside that global guarantee.

  • Scenario-coverage structure: ARC-SC's scenario-coverage objective is normalized, monotone, and submodular for fixed response values in [0, 1].The structural result does not require predictive scenarios to be exact samples from a correlated Gaussian-process posterior.
  • Scenario-coverage structure: ARC-SC discounts candidates that repeat predictive-scenario coverage, assigning greater marginal value to responses in scenarios weakly represented by the current batch.Once a scenario is strongly covered, its uncovered mass is small and similar candidates provide little additional gain.
  • Scenario-coverage structure: The overlap term is large when two candidates respond strongly in the same predictive scenarios, inducing complementarity without explicit pairwise geometric repulsion.This provides a scenario-space mechanism for reducing batch redundancy.
  • Anchored completion: For fixed anchors, the residual coverage objective remains normalized, monotone, and submodular with respect to the completion set.Adding the same anchor set preserves monotonicity and diminishing marginal returns, while subtracting its constant coverage establishes normalization.
  • Anchored completion: Greedy scenario completion inherits the classical cardinality-constrained submodular approximation guarantee when anchors and the admissible completion pool are fixed.The guarantee applies to the remaining B − K positions and approximates the optimal residual scenario coverage over that same pool.
  • Scope of the guarantee: The guarantee does not extend globally to the implemented algorithm because its admissible set changes with the partial batch under the dynamic risk-support constraint.A penalized fallback is used when no risk-feasible extension exists.

4 EXPERIMENTS AND RESULTS

ARC-SC is evaluated against POF on frozen-oracle superconductivity and JARVIS benchmarks under controlled invalid-output mechanisms. It improves first-hit timing on superconductivity, while JARVIS results are directionally favorable but statistically inconclusive.

  • Experimental setup: The evaluation uses two frozen-oracle materials benchmarks, four invalid-output mechanisms, three invalid-rate levels, three initialization conditions, paired seeds, and five rounds of batch size five.Superconductivity targets high critical temperature; JARVIS targets low formation energy per atom, with response signs adjusted so larger values are preferred.
  • Evaluation and statistical analysis: The primary endpoint is the first closed-loop round containing a newly queried valid target hit, with no-hit runs censored at six.Uncertainty is assessed using an exact two-sided sign test and hierarchical-bootstrap confidence intervals based on 10,000 replicates.
  • Overall first-hit performance: 70/251/39 first-hit W/T/L outcomes favor ARC-SC on superconductivity, with conditional win fraction 0.642, sign-test p = 0.004, and Δeτhit = −0.269 [−0.539, −0.025].Both the win–loss imbalance and mean paired censored first-hit effect favor ARC-SC.
  • Overall first-hit performance: ARC-SC finds superconductivity targets earlier through both discovery-before-budget and earlier-round advantages, while reducing invalid evaluations by 5.9 percentage points on average.Cumulative valid hits increase by 0.283 per run, but the confidence interval includes zero, making first-hit timing the clearest supported benefit.
  • Overall first-hit performance: On JARVIS, 49/272/39 W/T/L outcomes and Δeτhit = −0.117 [−0.297, 0.053] are directionally favorable but statistically inconclusive.ARC-SC produces 18 additional aggregate valid target hits, while secondary-effect intervals include zero; the result is interpreted as robustness rather than established superiority.
  • Discovery probability across experimental rounds: Superconductivity first-hit probabilities are 15.3% versus 10.0% at round 1 and 35.6% versus 30.6% at round 5 for ARC-SC and POF, respectively.On JARVIS, the round-5 probabilities are 48.1% and 44.4%, with confidence intervals including zero.
  • Failure-mechanism analysis: Scenario coverage is most advantageous under superconductivity feature-interaction and low-validity failures, whereas several hybrid and off-manifold configurations produce ties.Mechanism-level results are descriptive rather than individually powered hypothesis tests.

5 CONCLUSION

The paper concludes that ARC-SC combines marginal valid-hit exploitation with risk-constrained scenario diversification for first-hit batch inverse design. Results support it as a context-dependent diversification strategy, with future work focused on adaptive diversification and candidate reachability.

  • Conclusion: ARC-SC preserves risk-admissible high marginal valid-hit candidates as anchors, then fills the batch with complementary predictive target scenarios.This construction is intended to reduce redundancy under predictive uncertainty while retaining marginal exploitation.
  • Conclusion: The scenario-coverage objective is normalized, monotone, and submodular, giving greedy completion the classical cardinality-constrained approximation guarantee for fixed anchors and a fixed admissible pool.The full implementation differs because its admissible set changes dynamically with the partial batch through the risk constraint.
  • Conclusion: Holdout-seed experiments show statistically supported first-hit improvement and fewer invalid evaluations on superconductivity, while ARC-SC remains competitive on higher-dimensional JARVIS.The results support ARC-SC as a POF-preserving diversification strategy rather than a uniformly superior replacement for marginal acquisition.
  • Future directions: Future work includes adaptive anchor counts and risk budgets, more faithful posterior scenarios, and joint treatment of target uncertainty, validity, support, and batch composition.The conclusion also identifies candidate generation and acquisition as coupled when unsupported high-dimensional designs limit candidate reachability.

A.1 IMPLEMENTATION AND REPRODUCIBILITY

The implementation uses shared predictive infrastructure for ARC-SC and POF, while POF selects the B largest marginal valid-hit scores.

  • Implementation and reproducibility: ARC-SC and POF share the target surrogate, validity model, observed data, target threshold, and, in uniform-pool experiments, the same finite candidate pool.The shared infrastructure isolates acquisition-strategy differences in paired comparisons.
  • Implementation and reproducibility: POF selects the B largest marginal valid-hit scores h(x) = p_y(x)p_v(x).This defines the ordinary top-B marginal baseline used in the experiments.

A.1.1 UNIFORM CANDIDATE-POOL CONSTRUCTION

Uniform candidate pools are sampled from independently padded descriptor bounds and reused unchanged by ARC-SC and POF within each paired case.

  • Uniform candidate-pool construction: Each descriptor’s search interval is padded by ρ_b = 0.05 beyond its observed minimum and maximum.The interval is [ℓ_j − ρ_b(u_j − ℓ_j), u_j + ρ_b(u_j − ℓ_j)].
  • Uniform candidate-pool construction: At every round, N = 10,000 candidates are sampled independently and uniformly from the axis-aligned box.The protocol is used for all superconductivity experiments and selected JARVIS mechanisms.
  • Uniform candidate-pool construction: Paired uniform-pool cases supply the same condition- and round-specific candidate pool to ARC-SC and POF.This ensures the methods are compared on exactly the same candidates.

A.1.2 CLOSED-LOOP JARVIS TRUST-REGION CANDIDATE POOL

The JARVIS trust-region pool generates candidates from currently observed designs to avoid unsupported regions in challenging failure settings. Its trajectory-dependent construction uses only information available to the current method and can differ between ARC-SC and POF.

  • Global uniform sampling is dominated by unsupported candidates in 15-dimensional JARVIS space under off-manifold and hybrid failures.
  • Candidate anchors are sampled from currently observed valid rows when available, or from all observed rows otherwise.Descriptor-wise perturbations use observed-design interquartile ranges, with benchmark-reference fallbacks for degenerate scales.
  • Generated candidates are clipped to benchmark-reference feature bounds and duplicate rows are removed.
  • Trust-region pools use only current method observations and validity outcomes, excluding held-out labels, unqueried oracle values, future observations, and competitor trajectories.Consequently, the pools are trajectory dependent and need not match between ARC-SC and POF.

A.1.3 TARGET AND VALIDITY MODELS

The closed-loop models refit target and validity predictions each round using distinct observed-data subsets. ARC-SC and POF share the same model classes and seeds within each paired case.

  • The target surrogate is refitted each round using only valid observations and a standardized-input Gaussian-process regressor.Predictive means and standard deviations are denoted by µ(x) and σy(x).
  • The validity model uses all observed valid/invalid labels in a 300-tree random-forest classifier.If only one validity class is observed, a constant-probability model is used until both classes appear.
  • ARC-SC and POF use identical target-surrogate and validity-model classes with matched model seeds within each paired case.

A.1.4 EMPIRICAL-SUPPORT SCORE

The empirical-support score measures relative local support around each candidate rather than calibrated probability. It is recomputed from the current observed design geometry each round.

  • Candidates are standardized using current observed-design statistics and evaluated against their five nearest observed designs.
  • The distance scale is recomputed each round as the empirical 75th percentile of five-neighbor distances, with a lower bound of 10^-6.
  • The support score is a relative empirical-support score, not a calibrated probability.

A.1.5 PREDICTIVE SCENARIOS, RANK GATES, AND RISK FALLBACK

ARC-SC combines predictive scenarios, rank-normalized gates, and risk-aware batch completion, while its greedy guarantee applies only to fixed anchors and a fixed admissible completion pool. The benchmark evaluation uses frozen oracles constructed from the reference datasets.

  • Predictive scenarios, rank gates, and risk fallback: Independent marginal predictive draws preserve candidate-wise means and variances but omit cross-candidate posterior covariance.Correlated Gaussian-process sampling is enabled only for pools of at most 256 candidates; the reported N = 10,000 experiments therefore use independent draws.
  • Predictive scenarios, rank gates, and risk fallback: Rank gates transform marginal valid-hit, validity, and support scores using rank percentiles and specified exponents.The gates use ϵh = ϵv = ϵs = 0.05, αh = 1, and αv = αs = 0.5.
  • Predictive scenarios, rank gates, and risk fallback: ARC-SC selects risk-admissible anchors by h(x) and completes batches by maximizing scenario-completion gain over feasible extensions.When no feasible extension exists, it uses a penalized fallback with λR = 10.
  • Predictive scenarios, rank gates, and risk fallback: The risk budget is enforced whenever a feasible extension exists, rather than as an unconditional terminal-batch guarantee.
  • Predictive scenarios, rank gates, and risk fallback: The greedy completion guarantee compares against the optimal completion only for the same fixed anchors and fixed admissible pool.
  • Predictive scenarios, rank gates, and risk fallback: Frozen-oracle evaluation keeps oracle predictions fixed for both reference and newly generated candidates throughout closed-loop acquisition.Original dataset labels are used for oracle construction and validation, not alternated with oracle predictions during evaluation.
Loading 2608.27873v1…