Source-linked AI summary

Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery

Niranjan Srinivas, Debajyoti Ray, Elias Nakouzi

arXiv:2609.09413v1cs.AIcs.CEcs.RO

TL;DR

The paper asks how to choose recovery experiments that support scale-up decisions connecting laboratory results with product requirements, costs, and scale effects. It analyzes CICERO records with fitted-model active-learning benchmarks and proposes selecting batches by expected downstream Bayes-risk reduction. Adaptive policies reached the recorded NdFeB enrichment maximum in 16–24 wells versus 48 for space filling, while the authors identify data, economic, and scale-transfer requirements for prospective validation.

  • Problem

    Scale-up recovery decisions require laboratory results to be connected with product requirements, process costs, and intended-scale effects.

  • Method

    The paper analyzes CICERO records, reconstructs fitted-model adaptive searches, and proposes selecting batches by expected reduction in downstream Bayes risk.

  • Results

    Adaptive policies reached the recorded NdFeB enrichment maximum by 16–24 wells, compared with 48 for space filling.

  • Takeaways & Limitations

    Selecting an economically preferred process requires an explicit downstream loss and the inputs needed to evaluate it.

  • Takeaways & Limitations

    The archive cannot directly validate policies selecting unrecorded conditions and lacks a validated downstream loss and sufficient data for plant-scale transfer.

Abstract

from arXiv · show

Choosing a recovery process for scale-up requires connecting laboratory results with product requirements, process costs, and scale effects. We analyze records from Pacific Northwest National Laboratory's Computer Intelligence for Critical Element Recovery and Optimization (CICERO) workflow for autonomous selective precipitation. Active learning uses prior results to choose experiments. In a conditional retrospective benchmark with fitted models and recycled neodymium-iron-boron (NdFeB) magnet records, active learning finds the best recorded result with fewer experiments than nonadaptive space filling. Enrichment is the selected rare-earth-to-iron ratio relative to that in the feed. Adaptive policies reach the recorded enrichment maximum by 16 to 24 wells (individual experiments), versus 48. Our two-stage reconstruction ties two adaptive alternatives at 16 wells. Conditional analyses of recycled samarium-cobalt (SmCo) magnets show a Round 2 tradeoff between purity and nominal yield, the recovery fraction calculated from an assumed starting amount - NdFeB Round 1 routes differ in enrichment. Rankings for produced water from oil and gas extraction depend on phase and dilution assumptions requiring confirmation. We propose choosing batches by their expected reduction in downstream Bayes risk: the minimum expected loss among available process decisions under current beliefs. In exploratory simulations, a hybrid that filters candidates has lower estimated loss than the implemented joint search across routes and conditions. Differences involving the synthetic two-stage policy are small relative to estimation uncertainty. We outline a pre-registered prospective test under a shared loss and logging standard, requiring clarified measurements and records, a defined process decision and relevant outputs, credible economic inputs, and validation at the intended scale.

1 Background and motivation

Selecting recovery experiments for scale-up must account for product requirements, process costs, and scale transfer, not laboratory purity or yield alone. The paper frames adaptive experimentation around the eventual process choice and its downstream loss.

  • Scale-up recovery choices depend on product requirements, process costs, and how laboratory results transfer to the intended scale.
  • CICERO links feedstock characterization, technoeconomic reasoning, experimental planning, robotic execution, ICP-MS, and Bayesian optimization for selective precipitation.
  • Earlier work selected recovery pathways and then optimized conditions with different utilities, without a common experiment-selection criterion.
  • The CICERO campaigns considered did not treat experimental cost as material to route or condition selection, whereas this paper focuses on downstream process choice.
  • The proposed objective selects batches by their expected reduction in downstream Bayes risk across discrete routes and continuous conditions.
  • The workflow reviews records and benchmarks search on the recorded NdFeB pool, using synthetic comparisons to inform a prospective test.

2 What the current evidence supports

The records support conditional, assumption-dependent comparisons of SmCo, NdFeB, and produced-water recovery experiments, but not definitive process or scale-transfer choices. SmCo shows a purity–nominal-yield tradeoff, NdFeB shows route-specific enrichment contrasts, and produced-water rankings change with measurement interpretation.

  • SmCo: purity and nominal yield: Four Round 1 and eight Round 2 conditions define the recorded SmCo purity–yield frontiers.
  • SmCo: purity and nominal yield: At g = 0.95, the loss-minimizing recorded condition switches from H11 to B12 at ρ = 0.259.H11 has 113.86% nominal yield and 41.82% purity; B12 has 84.36% nominal yield and 96.92% purity.
  • NdFeB Round 1: conditional route contrast: 135, 43, 5.7, 3.3, 2.3, and 1.9 are hydroxide median enrichment values, versus 73, 240, 138, 255, 201, and 129 for oxalate.The within-route frontiers retain F01 for hydroxide and E07, D08, and H10 for oxalate; oxalate interpolation has log-RMSE 1.606.
  • Produced water: replication and interpretation: Direct Mg medians rank B, C, D, E, F, A, G, H, but normalization can select G or H depending on phase assumptions.Under condition-comparable signals, omitting any one well leaves the direct order unchanged; product-phase normalization selects G, while residual-phase readings select H.

3 Decision-focused formulation and algorithmic direction

The paper formulates experiment selection around downstream process decisions, defining batch value as expected reduction in downstream Bayes risk. It then outlines two-stage, joint, and hybrid policies that approximate this objective under route and condition constraints.

  • 3.1 Downstream loss: A deployment action combines a recovery route, operating condition, and experimental scale or assay fidelity.Technical uncertainty includes response, recovery, impurity, noise, and scale discrepancy; economic uncertainty includes product value, reagent price, disposal cost, and throughput value.
  • 3.1 Downstream loss: Deployment loss combines process cost and penalty at the intended target scale.The loss is written as L_s⋆(d, θ, ϕ) = C_process(d, θ, ϕ; s⋆) + C_penalty(d, θ, ϕ; s⋆).
  • 3.1 Downstream loss: Current Bayes risk is the minimum expected downstream loss among available deployment decisions under current technical and economic beliefs.The formulation can factor technical and economic beliefs, while allowing a joint distribution when dependence matters.
  • 3.2 Value of a batch: A feasible batch is valued by its expected reduction in Bayes risk after observing future outcomes.Under exact Bayesian updating and optimization, the value of information is nonnegative because the decision maker can ignore an observation.
  • 3.2 Value of a batch: With equal plate capacity and cost, batches are ranked by decision value; unequal fidelities, batch sizes, turnaround, or stopping require explicit cost treatment.Cost can enter as a budget constraint or be subtracted from decision value after conversion to common loss units.
  • 3.3 Candidate method family: Three candidate families are proposed: downstream-aware two-stage selection, direct joint decision-value search, and a structured hybrid.The two-stage policy retains route discrimination followed by within-route optimization; the joint policy optimizes approximate Δ(B) over route-condition actions; the hybrid allocates a finite batch by approximate downstream value after structured candidate generation.
  • 3.3 Candidate method family: The hybrid uses EC2-style route discrimination and route-specific GP criteria to generate candidates before allocating wells across routes, conditions, replicates, and controls.It samples plausible scenarios, identifies preferred routes, targets promising or decision-sensitive regions, and applies approximate marginal Δ.

4 Retrospective analyses and process decision model

The retrospective analyses show that adaptive search can find the recorded NdFeB enrichment maximum faster than tested nonadaptive designs, while emphasizing that the benchmark does not evaluate downstream-aware acquisition or calibrated uncertainty reduction. Process choices also require measurements, economic inputs, and scale-transfer assumptions.

  • 4.1 Valid retrospective evaluation: The produced-water plate contains 12 recorded wells at each of eight conditions, while magnet records generally provide one outcome per condition.These data support reveal-only completed-grid benchmarks, not repeated noisy trajectories or claims about the original grid-generating policy.
  • 4.1 Valid retrospective evaluation: Recorded-grid benchmarks select stored conditions and measure wells needed to find the exact maximum among fixed recorded E values.GP predictions guide selection, but stored outcomes determine performance; this differs from interpolation diagnostics.
  • 4.2 Conditional NdFeB finite-pool benchmark: 16 wells suffice for H10 under the two-stage reconstruction, routewise GP-UCB, and equal route split; mixed-route BO requires 24 and deterministic space filling 48.Random sampling remains 21.36 below the pool-optimum enrichment of 508.97 on average after 48 wells.
  • 4.2 Conditional NdFeB finite-pool benchmark: The first three adaptive policies tie throughout the benchmark, with a common initial batch revealing H12 at E = 493.11 near the pool maximum 508.97.Space filling also reveals H12 initially, so differing initial designs prevent isolating adaptation alone.
  • 4.2 Conditional NdFeB finite-pool benchmark: The benchmark does not test downstream-aware selection because terminal-only symbolic shortfall cannot distinguish conditions already meeting the fixed target.Once H10 is revealed, best-revealed enrichment is saturated despite different oxalate allocations.
  • 4.2 Conditional NdFeB finite-pool benchmark: At 16 wells, nominal 90% plug-in latent intervals cover 42.5% of 80 hidden recorded conditions under both two-stage and routewise models.The fitted-hyperparameter trajectories therefore do not establish calibrated uncertainty reduction.
  • Process decision model: Decision-focused evaluation requires laboratory-resolvable outputs, credible economic inputs, and scale-transfer quantities validated at corresponding scales.Relevant inputs include protocol, phase, units, dilution, recovery, grade, replicates, product specifications, costs, throughput, mixing, mass transfer, residence time, geometry, and robustness.

5 Exploratory synthetic design probe

Exploratory synthetic comparisons under a common normalized downstream loss favor a structured hybrid proxy over direct joint decision-value search, while differences involving synthetic two-stage are small relative to estimation uncertainty. The probe uses a fixed factorial design, common seeds, and finite acquisition and terminal particle budgets.

  • Policy comparison: The comparison evaluates synthetic two-stage, joint decision-value, and structured hybrid proxy policies under a common normalized downstream loss.The hybrid proxy filters promising or uncertain candidates before applying the joint policy’s decision-value rule and is not the proposed EC2/GP-UCB hybrid.
  • Experimental design: The 32-cell factorial varies route gap, batch capacity, scale discrepancy, loss shape, and grade-specification uncertainty.Each cell uses 100 world seeds, resource budget 16, laboratory and bench costs of one and two, 256 acquisition particles, and 4,096 terminal particles.
  • Results: Mean excess losses are 0.041 for the structured hybrid proxy, 0.045 for synthetic two-stage, and 0.046 for the joint decision-value policy.Excess loss is chosen-action loss minus synthetic-oracle loss, with not deploying assigned loss one.
  • Results: The largest policy contrast is roughly a tenth of the policies’ mean excess loss.Reported standard errors are conditional on the configured synthetic model and acquisition setup.
  • Results: The structured hybrid proxy has lower estimated loss than direct joint decision-value search, while contrasts involving synthetic two-stage are small relative to world-seed Monte Carlo standard errors.The aggregate pattern remains stable across 64, 128, and 256 acquisition particles, but individual acquisition paths do not.

6 Proposed prospective campaign

The proposed prospective campaign is a preregistered, multi-batch comparison of decision-focused search under a shared loss and logging standard. It requires credible routes, explicit measurements and constraints, common evaluation, and confirmation at bench or pilot scale.

  • Campaign design: A prospective campaign should use at least two credible routes with continuous within-route conditions and fit two or three adaptive batches.The campaign should be small enough to permit adaptive batching but large enough that exhaustive testing is impractical.
  • Preregistration: Before experimentation, preregister routes, condition bounds, outputs, downstream loss or economic scenarios, candidate pool, constraints, budget, controls, stopping rule, and metrics.An unchanged preferred decision should be explicitly treated as a valid result.
  • Adaptive batches: The first batch should cover routes, feasibility, and the downstream decision boundary; the second should compare incumbent and decision-focused allocations using the same information.Where practical, both recommendations should share a plate, and bench or pilot confirmation should include the selected condition plus one credible alternative.
  • Logging standard: Each round should retain candidate sets, exclusions, code and model state, acquisition values, executed batches, manual changes, timestamps, resource use, raw measurements, dilution, QC metadata, and stopping decisions.Versioned pipelines and interoperable schemas should convert these records into decision-ready variables.

7 Related work, limitations, and open questions

The paper situates its decision-focused objective within Bayesian experimental design and identifies archive, measurement, provenance, economic, and scale-transfer limits that constrain interpretation.

  • Related work: Bayesian experimental design values observations by their expected effect on later decisions, motivating expected downstream Bayes-risk reduction as the paper’s objective.The related work connects this objective to knowledge-gradient methods and to a companion two-stage route-selection and within-route optimization structure.
  • Limitations: The CICERO archive cannot validate policies selecting unrecorded conditions because missing outcomes remain unavailable, even with better metadata.Most conditions also lack independent replicates, limiting repeatability estimates.
  • Limitations: Policy trajectories are unavailable, and changes between rounds may combine chemistry and protocol effects, limiting reconstruction of the original campaign.These constraints affect interpretation of retrospective comparisons across rounds.
  • Open questions: SmCo nominal-yield interpretation requires confirming whether the 3928.381 mg/L feed concentration applies to both rounds and how yields above 100% should be treated.Round 2 selection reconstruction also requires unavailable Bayesian-optimization recommendations and their mapping to executed wells.
  • Open questions: NdFeB route comparisons require executed maps, reagent concentrations, measured phases, dilution treatment, and the intended separation-factor definition.Cross-round recovery additionally requires the enriched-stock preparation and recovery denominator.
  • Open questions: Produced-water rankings depend on the executed map, stock concentration, signal units, dilution treatment, and measured phase, while replication requires clarifying column roles and outlier treatment.Linking plate row B to the bench experiment also requires confirming whether it informed that experiment and whether the feed was comparable.
  • Open questions: An economically preferred process requires the Sm product-grade threshold, rework cost, recovered-Sm value, and decision inputs beyond NdFeB enrichment alone.The enrichment score omits recovery, reagent demand, throughput, waste, and downstream purification.

8 Conclusion

The conditional benchmark finds adaptive policies reach the recorded NdFeB enrichment maximum sooner than space filling, while the paper proposes a prospective test based on downstream decision loss.

  • Conclusion: 16–24 wells versus 48: adaptive policies reach the recorded NdFeB enrichment maximum sooner than space filling in the conditional MODEL-BASED benchmark.The two-stage reconstruction shares the earliest 16-well discovery budget with two adaptive alternatives.
  • Conclusion: Conditional analyses identify SmCo purity–nominal-yield tradeoffs, differing NdFeB route enrichment profiles, and produced-water rankings dependent on measurement interpretation.Selecting an economically preferred process requires an explicit downstream loss and the inputs needed to evaluate it.
  • Conclusion: Expected reduction in downstream Bayes risk is proposed for selecting experimental batches.The proposed objective targets the expected reduction in the loss of the eventual process choice.
  • Conclusion: A prospective comparison would use a common downstream loss and logging standard after resolving archive gaps and defining the deployment decision, outputs, and credible economic ranges.The proposed test is intended to evaluate batch selection under clarified decision inputs and records.

A Finite-pool policy table

The finite-pool table reports reveal-only NdFeB grid simple regret by well budget, with random designs averaged over seeded runs and other policies deterministic.

  • Finite-pool policy table: Grid simple regret measures enrichment shortfall from the recorded-pool optimum using the best revealed condition.Random-design entries average 20 seeded runs, while other policy entries are deterministic.
  • Finite-pool policy table: Each policy’s permitted terminal action defines symbolic downstream metrics, with the two-stage action remaining within its committed route.This differs from grid simple regret, which uses the best revealed condition.

B Robustness and descriptive analyses

Robustness and descriptive analyses show sensitivity to modeling choices, while several reported tradeoffs remain stable under specified alternative treatments.

  • Produced water: Produced-water Mg-only leaders and Mg–Ca frontiers are unchanged under retain-all, Tukey 1.5-IQR, and modified-z 3.5 sensitivity rules.The main analysis retains all measurements.
  • NdFeB Round 2: No cross-round NdFeB recovery or material balance is calculated because the enriched-stock denominator is unavailable.Within-plate profiles retain all 96 recorded values, including finite negative signals.
  • Finite-pool sensitivity: 16–24 wells: two-stage and routewise policies first discover the grid optimum across fitted and four perturbed GP specifications.Equal route splitting spans 16–40 wells, while mixed-route Bayesian optimization spans 24–40 wells; these scenarios probe sensitivity rather than provide uncertainty intervals.
  • Synthetic particle sensitivity: A structured hybrid proxy has lower point-estimated loss than joint decision-value search by −0.0052, −0.0035, and −0.0043 at 64, 128, and 256 particles.Negative values favor the first-named policy, and the estimates are point estimates only.
Loading 2609.09413v1…