Source-linked AI summary

Text-guided flow matching enables sample-efficient crystal structure generation

Wentao Li

arXiv:2609.01076v1cond-mat.mtrl-scics.AI

TL;DR

Crystal-generation controls often use algorithmically convenient variables rather than expressive combinations of realistic materials constraints. TFMat adds structured-text conditioning to CrystalFlow, and improves one-sample CSP match rates while aligning de novo population statistics, though broader validation is still required.

  • Problem

    Crystal-generation control variables are often poorly matched to realistic constraints, which rarely arrive as isolated labels.

  • Method

    TFMat extends CrystalFlow with a condition vector from frozen MatSciBERT embeddings of structured prompts containing composition, symmetry descriptors, and selected properties.

  • Results

    TFMat improves one-sample CSP match rates across Perov-5, Carbon-24, and MP-20, while de novo outputs show improved marginal distribution alignment and coarse property consistency in composition-selected candidates.

  • Takeaways & Limitations

    Structured materials text can serve as an actionable, inspectable prior over a flow-matching crystal generator’s search trajectory.

  • Takeaways & Limitations

    Density-functional relaxation and unfiltered-population analysis are required before claims about thermodynamic stability of newly generated materials.

Abstract

from arXiv · show

Crystal generators can now propose periodic structures, but their control interfaces remain poorly matched to the mixed descriptors used in materials design. Text provides a compact way to combine composition, symmetry, prototype and property cues, yet it has not been clear whether such information can steer flow-based crystal generation. Here we introduce TFMat, a text-conditioned flow-matching framework that uses structured materials language as a semantic prior for a CrystalFlow generator. Across Perov-5, Carbon-24 and MP-20 crystal structure prediction benchmarks, TFMat improves one-candidate match rates over CrystalFlow and reaches a 92.04% MP-20 match rate with 20 candidates; in de novo generation, it improves element-count and density distribution alignment while retaining coarse property consistency in composition-selected outputs. These results position structured text as an inspectable control layer for translating human-readable materials intent into candidate crystals for downstream simulation and validation.

Introduction

Crystal generation has advanced from prototype enumeration toward constrained periodic-structure design, but current controls often remain too narrow for mixed materials intent. TFMat addresses this gap by adding structured text conditioning to CrystalFlow while preserving its geometric flow model.

  • Motivation: Generative crystal models search compositions, lattices, and atomic arrangements rather than only enumerating known prototypes.The design space is enormous, while stable crystals occupy a sparse subset.
  • Prior approaches: Recent graph, adversarial, diffusion, and flow-based methods improved geometric modeling of periodic coordinates, lattices, and atom types.These approaches established that chemical and symmetry constraints can improve crystal validity.
  • Control gap: Materials constraints are often mixed across formulas, prototypes, symmetry, stability ranges, and electronic-property expectations rather than supplied as isolated labels.This makes scalar or composition-only conditioning a narrow representation of materials intent.
  • TFMat: TFMat adds frozen MatSciBERT embeddings of structured materials prompts to CrystalFlow without replacing the geometric backbone.Its prompts include formula or composition, symmetry descriptors, and selected scalar properties, excluding lattice vectors and fractional coordinates.
  • Evaluation: TFMat is evaluated in composition-fixed crystal structure prediction and text-conditioned de novo generation across Perov-5, Carbon-24, and MP-20.The framework improves one-sample match rates over CrystalFlow, with the largest gains on constrained perovskite and carbon-network benchmarks.

Results

TFMat improves structural retrieval in CSP and reshapes population-level statistics in DNG, while preserving geometric plausibility and showing coarse property consistency in a composition-selected subset.

  • Crystal structure prediction: TFMat increases one-sample CSP match rates over CrystalFlow across Perov-5, Carbon-24 and MP-20.Match rates rise from 53.69% to 91.33%, 15.02% to 46.40%, and 67.65% to 77.80%, respectively.
  • Crystal structure prediction: 92.04% MP-20 match rate with 20 candidates makes TFMat the strongest compared method for repeated-sampling retrieval.CrystalFlow reaches 85.38% and TGDMat (Long) 82.02%; TGDMat variants achieve lower RMSE, separating retrieval from refinement.
  • De novo generation: TFMat lowers MP-20 DNG element-count EMD from 0.2489 to 0.1990 and density EMD from 0.1701 to 0.0794 relative to CrystalFlow.Structural validity remains 99.06% and coverage precision 99.67%, while atom-type generation remains the main bottleneck.
  • Property consistency: In a composition-selected subset, CGCNN diagnostics show 95.94% formation-energy sign agreement and 83.65% band-gap zero/nonzero agreement.These are surrogate diagnostics for selected outputs, not evidence of DFT stability, synthesizability or unfiltered population-wide property control.
  • Population diagnostics: The text-conditioned population shows modestly improved distributional alignment, with per-panel t-SNE Jensen-Shannon divergence decreasing from 0.6367 to 0.6200.A separate shared-coordinate diagnostic reports the same direction, but independently projected t-SNE panels should not be compared by point coordinates.

Discussion

TFMat’s structured text conditioning improves sample-efficient structural retrieval while preserving CrystalFlow’s geometric generator, but its evidence remains bounded by prompt scope, surrogate-based property analysis, and incomplete final refinement. The results support using embedded materials descriptions as an inspectable control layer, not unrestricted natural-language discovery.

  • Results: TFMat improves match rate across Perov-5, Carbon-24, and MP-20 in the one-sample CSP regime, where semantic priors are most visible.The model retains CrystalFlow as the central geometric generator while improving the chance of entering a plausible structural basin.
  • Comparison: TFMat is complementary rather than a universal replacement for diffusion-based text guidance, because TGDMat remains strong for RMSE and MP-20 DNG compositional validity.The comparison frames TFMat’s contribution as sample-efficient structural retrieval with continuous flow matching.
  • Distributional effects: DNG and diagnostic results show changes in element-count and density distributions, graph-embedding overlap, and composition-selected property consistency beyond validity alone.The property analysis relies on CGCNN surrogates and a selected subset, while compositional validity trails the best published text-guided diffusion result.
  • Scope: The prompts demonstrate metadata-conditioned generation rather than independent discovery from unconstrained natural language.They use database-derived formula or composition information, space-group descriptors, and scalar properties; broader ablations and out-of-distribution checks remain necessary.
  • Validation: Density-functional relaxation and unfiltered-population analysis are required before claims about thermodynamic stability of newly generated materials.These checks define a practical boundary on interpreting the generated candidates as stable materials.
  • Implications: Embedding structured descriptions into a geometric flow makes text an actionable, inspectable prior for future closed-loop crystal-discovery workflows.The interface can be inspected, edited, and reused by researchers and automated materials-science systems.

Methods

TFMat evaluates text-conditioned crystal generation across CSP and DNG settings using structured metadata embeddings with a CrystalFlow flow-matching backbone. The method predicts lattice, coordinates, and, for DNG, atom types through learned velocity fields integrated with Euler updates.

  • Evaluation: TFMat evaluates Perov-5, Carbon-24 and MP-20 using benchmark splits for crystal structure prediction and de novo generation.CSP fixes atom types and composition, whereas DNG generates atom types, lattice and fractional coordinates jointly after sampling atom count.
  • Text conditioning: Structured prompts combine formula-level identity, symmetry information and selected scalar properties into a semantic condition rather than supplying raw coordinates.For MP-20, fields can include formula, crystal system or space group, formation energy, band gap and energy above hull.
  • Flow matching: The flow-matching model transports samples from a simple source distribution to crystal data through a learned velocity field.Crystal states use a six-dimensional lattice-polar representation and fractional coordinates, with periodic displacement handled by the minimum-image convention.
  • Velocity prediction: The network predicts lattice and coordinate velocities from the current state, fixed atom types, time and projected text condition.For CSP, the predicted velocities are trained against target velocities; DNG additionally predicts the relaxed atom-type representation.
  • Sampling: Sampling integrates the learned velocity field with Euler updates over T steps, using a default conditional sampling factor of 1.0.Fractional coordinates are updated modulo 1, while guidance-factor sweeps are treated as supplementary rather than as an isolated contribution.
  • Metrics: CSP match rate and RMSE use StructureMatcher tolerances, while DNG reports validity, coverage and Earth mover’s distances for population-level diagnostics.The evaluation also includes t-SNE embeddings and CGCNN-based formation-energy and band-gap diagnostics.

Supplementary Information

The supplementary analyses define the benchmark scope, structured-prompt interface, metric interpretation and control experiments. They report stronger one-sample CSP recovery and improved DNG distribution alignment, while documenting limits involving exact prompt recovery, surrogate filtering and t-SNE comparability.

  • Dataset description: Perov-5, Carbon-24 and MP-20 span constrained perovskites, larger carbon networks and inorganic structures with up to 20 atoms per unit cell.The MP-20 t-SNE diagnostic uses 9,046 held-out structures.
  • Evaluation regimes: CSP fixes composition and predicts lattice and fractional coordinates, whereas DNG jointly generates atom types, lattice and coordinates after sampling atom count.These regimes use the same data resources but impose different generation constraints.
  • Model and prompts: TFMat retains CrystalFlow’s flow-matching generator and adds a projected structured-language condition, isolating the conditioning pathway as the main architectural change.Prompts are assembled reproducibly from database fields such as composition, prototype, symmetry and selected properties, then encoded with frozen MatSciBERT and mean pooled.
  • Metric interpretation: CSP match rate is reference-based, whereas DNG validity, coverage and distribution metrics do not establish exact prompt recovery, thermodynamic stability or synthesizability.The metric split motivates reporting several diagnostics rather than a single score.
  • CSP results: Under 20 evaluations, TFMat remains strongest on MP-20, while Perov-5 and Carbon-24 show lower RMSE but slightly lower match rate than CrystalFlow.This establishes a task-dependent tradeoff rather than uniform dominance across metrics.
  • Prompt scope: Structured prompts are reproducible metadata interfaces, but they are weaker than coordinate conditioning because they omit fractional coordinates and lattice vectors.Field availability varies across datasets and records, so every prompt need not contain every field.
  • Prompt controls: Property-shuffle controls retain most of the full-text gain when formula and symmetry remain fixed, consistent with those fields being dominant CSP cues.Partial embedding perturbations provide an additional prompt-control analysis under one-candidate evaluation.
Loading 2609.01076v1…