Source-linked AI summary
Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening
Zhilong Song, Lixue Cheng
TL;DR
Crystal generators produce candidates faster than DFT and experiments can assess them, while existing screens often test little beyond atomic overlap and lack chemical failure explanations. The paper uses autonomous agents and active refutation to discover eight PRIS laws, then shows that they diagnose damaged structures, track synthesizability, and reduce expensive validation while preserving target-reaching candidates.
Problem
Crystal generators and tool-using agents propose structures faster than DFT calculations or experiments can assess them, while common screens provide little chemical reason for failure.
Method
Autonomous agents generated, implemented, and actively refuted candidate laws from experimental structures, interpretable quantities, and prescribed crystal damage.
Results
Active refutation rejected eleven initially successful conclusions and left eight compact laws, while PRIS explains structural plausibility across screening, synthesizability, and crystallographic identity tests.
Takeaways & Limitations
PRIS turns structural screening into a chemical diagnosis in which each violation names the failing mechanism and can guide subsequent calculations and experiments.
Takeaways & Limitations
Small relaxation energy rules out gross strain but not the damage regimes examined in the study.
Abstract
from arXiv · showhide
Crystal generators and tool-using agents propose structures faster than density functional theory (DFT) energy and phonon calculations or experiments can assess them. Deciding which candidates merit expensive assessment is therefore the bottleneck, yet most screens test little beyond atomic overlap and give no chemical reason for failure. Here, our agents generate, test and actively refute two million candidate laws, leaving eight Plausibility Rules for Inorganic Structures (PRIS). These laws encode five mechanisms: short-range repulsion, ionic contact and packing, electrostatic balance, bond-valence conservation and crystallographic site complexity. Experimental structures satisfy our law sets at 82--99%, but satisfy Pauling's rules 2--5 together at only 6.5%. The strictest set detects 87.9% of damaged crystal structures, whereas distance cutoffs detect only 1.6--3.2%. PRIS plausibility is linearly correlated with synthesizability, so the PRIS-derived synthesis score (PSS) explainably screens 83.7% of hard-to-synthesize structures while retaining 80.7% of experimental structures. In a property-conditioned inverse-design run, PRIS and PSS can reduce the DFT validation queue by up to 67.3% and keep 99.2% of the candidates whose DFT-validated bulk moduli reach the design target. Beyond screening, PRIS explains why GNoME remains enriched in rare low-symmetry structures and reveals how wrong-element assignments in falsified crystal reports hide behind plausible coordinates. PRIS moves screening from a pass-or-fail verdict to a chemical reason for failure, showing that autonomous agents can discover, by active refutation, physicochemical laws that guide calculations and experiments.
1 Introduction
Crystal discovery now faces a screening bottleneck: generated structures arrive faster than DFT and experiments can assess them, while common screens provide little chemical diagnosis. The paper develops interpretable laws that retain plausible structures, detect damaged ones, and identify the failing physical or chemical constraint.
- 6.5% of charge-balanced ionic experimental structures satisfied Pauling’s rules 2–5 together, while 0.5- and 0.7-Å distance cutoffs detected only 1.6% and 3.2% of chemically damaged structures.The classical rules were too idealised when applied jointly, whereas distance cutoffs were element-blind and chemically weak.
- A useful structural law must retain experimental structures, detect damaged ones at high rates, and explain why a candidate is implausible.Existing screens do not establish what a criterion rejects, and distance cutoffs examine neither chemical ordering nor elemental identity.
- 2,037,606 candidate evaluations were actively refuted, leaving eight compact Plausibility Rules for Inorganic Structures (PRIS).Failed claims remained recorded, so progress was measured by refutation rather than proposal volume.
- PRIS encodes short-range repulsion, ionic contact and packing, electrostatic balance, bond-valence conservation, and crystallographic site complexity.Each unsatisfied law identifies the mechanism that fails.
- 82–99% experimental-structure satisfaction and 87.9% detection of chemically damaged structures show that PRIS combines retention with structural-error detection.No synthesis label entered PRIS selection, yet PRIS plausibility was linearly correlated with synthesizability.
- PSS screened 83.7% of hard-to-synthesize structures while keeping 80.7% of experimental ones, and PRIS/PSS reduced the DFT validation queue by up to 67.3%.The inverse-design screen retained 99.2% of structures reaching the bulk-modulus target under DFT, with every removal traceable to a named mechanism.
- PRIS traces GNoME’s low-symmetry excess to artificial ordering and exposes wrong-element assignments hidden behind plausible coordinates.These mechanisms extend screening beyond overlap checks to chemical ordering and elemental identity.
2 Results
Autonomous proposal, testing and refutation produced eight PRIS laws that diagnose structural failures through complementary physicochemical mechanisms. Across held-out validation and screening tasks, the laws and PSS improved damage detection, synthesizability screening, and DFT-queue prioritization while exposing trade-offs between strictness and retention.
- Autonomous law discovery: Two million candidate-law evaluations left eight one-line PRIS laws spanning five physicochemical mechanisms.Each law identifies an unsatisfied constraint and the mechanism to review.
- Held-out validation: Set 4 reached 81.8% held-out experimental-structure satisfaction and 91.1% overall damage detection, with at least 73.4% detection in every class.It combined laws targeting different failure modes rather than tightening one cutoff.
- Held-out validation: Only 6.5% of 5,297 held-out charge-balanced ionic experimental structures satisfied Pauling rules 2–5 jointly, whereas PRIS occupied the useful region between excessive rejection and weak detection.Distance cutoffs detected only 1.6–3.2% of damage, while Pauling’s rules rejected too many experimental structures.
- Synthesizability screening: Set 4 retained 80.7% experimental satisfaction while screening 51.9% of hard-to-synthesize structures, compared with 83.7% screened by PSS.PSS is continuous and serves as an operating point for adjusting queue strictness.
- Synthesizability screening: Mean Set 4 violation fell from 50.3% to 30.4% across CLscore deciles while mean PSS rose from −11.37 to −0.18, with R2 = 0.99 and 0.92 for the two models.The authors report a linear relation between plausibility and predicted synthesizability, while noting that PRIS does not order plausible polymorphs.
- Property-conditioned screening: PRIS and PSS reduced the DFT-validation queue by up to 67.3% while retaining 99.2% of candidates in the DFT-validated high-property subset.The retained high-property subset included all 140 candidates that Set 4 would remove.
3 Discussion
PRIS provides an independent structural-plausibility layer that diagnoses chemical failure across validation queues, databases, and individual sites. Its mechanisms remain separate, while PSS extends plausibility into synthesis-oriented ranking.
- 3 Discussion: PRIS operates at the validation-queue, database, and individual-site levels, adding structural plausibility before stability and synthesizability assessments.The check asks whether an arrangement obeys physicochemical laws and returns a falsifiable diagnosis.
- 3 Discussion: PRIS traces part of GNoME’s low-symmetry excess to artificial site splitting and exposes incorrect occupants at fixed coordinates.Coordinate-based screening cannot detect the latter failure.
- 3 Discussion: Each unsatisfied law names a failure mechanism, while Set 1–Set 4 demand fuller crystal models as task risk rises.The five mechanisms remain separate without an imposed hierarchy.
- 3 Discussion: PSS converts PRIS’s population-level link to predicted synthesizability into a continuous, tunable score for ranking candidates.The score weighs competing mechanisms, including dense packing and open packing with complete site splitting.
- 3 Discussion: PRIS and PSS remove chemically suspect candidates before expensive computation while linking each anomaly to a measurable physical quantity.The same verdicts can also label database entries and training data.
- 3 Discussion: Active refutation rejected eleven initially successful conclusions and left eight compact laws, with preserved records linking laws to counterexamples and checks.The workflow distinguished benchmark artefacts from real physics in an autonomous-agent process.
- 3 Discussion: The resulting screen goes beyond permissive distance cutoffs by explaining why structures are implausible and guiding subsequent calculations and experiments.This turns a binary verdict into a chemical diagnosis.
4.1 Autonomous-agent workflow, model and human oversight
The autonomous workflow coordinated agents that proposed hypotheses, ran analyses, wrote code, and diagnosed failed claims within human-defined goals and data boundaries. Human authors verified final outputs and retained responsibility for conclusions.
- 4.1 Autonomous-agent workflow, model and human oversight: The programme ran under one continuous analysis protocol from 27 July to 14 August, with procedures fixed before evaluation.Agent records, refuted conclusions, and reproduction commands were reported in supplementary sections.
- 4.1 Autonomous-agent workflow, model and human oversight: The agents reasoned by calling Codex through a custom client rather than making large language model calls independently.Codex used GPT-5.6-sol as its base model.
- 4.1 Autonomous-agent workflow, model and human oversight: Agents proposed hypotheses, designed and ran analyses, wrote code, constructed checks, and diagnosed failed claims with data access.The multi-agent loop coordinated task assignment, parallel execution, and result exchange.
- 4.1 Autonomous-agent workflow, model and human oversight: Humans set the initial goal, redirected study scope, and authorised access to reserved data without supplying a detailed scientific hypothesis.Confirmation analyses followed written procedures.
- 4.1 Autonomous-agent workflow, model and human oversight: Archived scripts and feature tables reproduce the reported values.
4.2 First-principles verification
First-principles verification re-derived the learned quantities and tested structural mechanisms with frozen plane-wave DFT protocols. The campaign evaluated energy, ordering, relaxation, and bulk-modulus behavior across 1,917 tasks.
- 4.2 First-principles verification: Four quantities were initially learned rather than computed and were re-derived with plane-wave DFT under a protocol frozen before job submission.The calculations used VASP 6.3.0 with PBE PAW potentials.
- 4.2 First-principles verification: The tests measured the energy landscape along ρ, ordering energies across symmetry-distinct merge groups, relaxation-energy release, and bulk moduli.Bulk moduli came from third-order Birch–Murnaghan fits to five-point energy–volume curves.
- 4.2 First-principles verification: Cells were fully relaxed before the relevant energy and bulk-modulus analyses.Machine-learning and DFT relaxation energies used matched per-cell definitions.
- 4.2 First-principles verification: The campaign comprised 1,917 tasks, with complete protocol, convergence, and fit-sensitivity analyses documented in Supplementary Section S20.Relaxed cells for 260 design candidates were provided as supplementary data.
4.3 Data sets and study design
The study used large experimental and damaged-structure datasets with structure-based partitions and held-out evaluation. A split-labelling error limited the independence of one reserve threshold fit.
- 4.3 Data sets and study design: Written protocols fixed allowed structural quantities, data splits, and success criteria before evaluation.Uncertainty was clustered by reduced composition, and damaged structures inherited their parent’s composition.
- 4.3 Data sets and study design: The dataset included 99,162 experimental structures from ICSD and COD, with structures assigned to discovery, held-out, or reserve partitions by seeded identifiers.Damaged structures inherited their experimental parent’s assignment.
- 4.3 Data sets and study design: The split separated structures rather than chemistries, so structures sharing a reduced composition could occupy different partitions.
- 4.3 Data sets and study design: Thresholds were fitted on 12,632 experimental and 8,590 damaged discovery structures, then assessed without refitting on 5,297 experimental and 3,612 damaged held-out structures.Variant tests evaluated transfer of added mechanisms rather than de novo discovery.
- 4.3 Data sets and study design: A split-labelling error allowed reserve structures into one full-sample threshold fit, so the reserve was not a fully independent final test.The incident was reported in Supplementary Section S19.
4.4 Structural descriptors and law-set evaluation
The evaluation computes eight law quantities under fixed conventions, applies explicit chemical and structural thresholds, and treats any evaluable violation as implausibility. Missing required inputs instead produce a separate no-verdict outcome.
- Conventions: Eight law quantities were computed under a fixed convention, with formal oxidation states assigned by integer charge balancing.External analyses could use a non-integer mean-valence fallback without refitting laws.
- Conventions: Native CIF oxidation annotations and bond-length-based valence inference were excluded to prevent bond lengths entering both charge assignment and bond-valence evaluation.
- Structural descriptors: Primary neighbours came from CrystalNN in pymatgen, while Law 7 used spglib with symprec=0.01.Discovery used deposited cells; external Law 7 comparisons used a common primitive-cell convention.
- Law evaluation: Each law combines a structural quantity with a directional threshold and, where needed, an explicit chemical condition.
- Law evaluation: Any violation among evaluable laws marks a structure implausible, whereas a missing charge, radius, or other required input produces no verdict.Only features available for more than 90% of experimental structures were admitted to the search.
4.5 Chemically damaged structures and law selection
The study evaluates law sets against composition-preserving chemically damaged structures and selects laws for damage detection subject to experimental-structure satisfaction. It then tests fixed and transferred thresholds against distance-cutoff benchmarks.
- Damage generation: Five composition- and stoichiometry-preserving perturbations generated damaged structures: uniaxial compression, cation–cation exchange, random displacement, isotropic expansion, and cation–anion exchange.Independent contact, electrostatic, and MatterSim-relaxation checks measured their physical and energetic severity.
- Law selection: Law selection maximised damage detection subject to an experimental-structure satisfaction floor.Fixed Set 4 was evaluated once on held-out data and compared with a provably optimal depth-three decision tree.
- Robustness: Threshold-transfer analysis re-derived each continuous cutoff at the corresponding held-out percentile and measured resulting verdict changes.Leave-one-damage-class-out evaluation reran tree and single-threshold selection in full.
- Robustness: Set 4 additions were reselected on a Set 3 base that had already seen all five damage classes, testing transfer of added mechanisms rather than de novo discovery.
- Benchmarking: The distance-cutoff benchmark compared PRIS with fixed 0.5- and 0.7-Å cutoffs and a cutoff matched to Set 4 satisfaction.The comparison used 440 experimental structures and 2,024 composition-preserving damaged variants.
4.6 External evaluation, synthesizability screening and statistics
External evaluation applies fixed PRIS thresholds across generated and deposited structures, while synthesizability and inverse-design analyses use independently trained models and held-out screening thresholds. The study also specifies sampling, ranking, data-release, and reproducibility procedures.
- External evaluation: Fixed PRIS thresholds were externally evaluated on GNoME, seven crystal generators, MatterSim-relaxed outputs, and historical falsified depositions.The protocol sampled 5,000 of 554,054 released GNoME structures and separately relaxed 500 raw MatterGen outputs.
- External evaluation: Thermodynamic stability, dynamical stability, and experimental record were compared across 26,600 Materials Project structures.
- Statistics: Same-composition ranking covered 18,920 pairs spanning 1,508 compositions, with compositions weighted equally and ties excluded from binary accuracy.Composition-cluster splits were used.
- Synthesizability screening: PSS was fitted as an antisymmetric, zero-intercept logistic score of six standardised descriptors, using frozen development-set medians for unavailable inputs.
- Synthesizability screening: Synthesizability trends were assessed with 50-bag CGCNN-PU and 50 MLP PU heads, whose consensus defined the hard-to-synthesize cohort.Training used 99,162 experimental and 8,125,976 unlabelled structures, plus frozen 128-dimensional MatterSim representations.
- Inverse design: Thirteen independently seeded MatterGen runs conditioned on a 400 GPa bulk modulus yielded 1,081 unique candidates for inverse design.Independent UMA predictions defined the high-property subset, and 541 experimental high-property structures fixed the screening threshold on two available PSS descriptors.
- Data and reproducibility: The public benchmark uses only Crystallography Open Database data, while derived scalar features, split assignments, numerical figure data, and analysis code are provided.The 260 DFT-relaxed design candidates are supplied as one CIF per candidate with accompanying indices.