Source-linked AI summary

Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks

Kai Gu, Haizheng Zhong

arXiv:2608.14734v1cs.LGcond-mat.mtrl-scics.AI

TL;DR

Deep-learning models predict nanocrystal properties but provide limited insight into the mechanisms determining size. This paper develops NanoEQL, a fully white-box neural network, and finds that nanocrystal size is described by a linear equation of three capability-related scalars.

  • Problem

    Deep learning’s black-box prediction process limits insight into the mechanisms governing nanocrystal synthesis.

  • Method

    The paper develops NanoEQL, a fully white-box neural network designed to unravel nanocrystal size-determination mechanisms.

  • Results

    Nanocrystal size is described by a linear equation composed of scalars for nanocrystallization, growth, and external input capabilities.

  • Takeaways & Limitations

    The resulting capability-related scalars provide an interpretable description of nanocrystal size determination.

  • Takeaways & Limitations

    Standard reciprocal, square-root, and cube-root operators exhibit excessively fast gradients near zero.

Abstract

from arXiv · show

Deep learning models of nanocrystal synthesis enable the prediction of size and shape by encoding precursors and reaction conditions. However, their black-box nature hinders gaining deep insights into the underlying synthetic mechanisms. Here, we develop the Nanocrystal Equation Learner (NanoEQL), a fully white-box neural network to unravel the size determination mechanisms of nanocrystal synthesis. Building on the EQL architecture, eight operators are introduced to replace standard activation functions to fit the mathematical equations in nanocrystal synthesis. Among these operators, three smoothed operators address the gradient explosion of singular operators at zero. To evaluate the weights of different precursors, we develop a temperature-gated attention pooling strategy that encodes concentration-driven and reactivity-driven chemical synthesis mechanisms into the temperature gate. The NanoEQL model illustrates that the final nanocrystal size can be described by a linear equation composed of three scalars representing nanocrystallization capability (-Zp), growth capability (Zrea), and external input potential (-Zops). These interpretable scalars not only advance the rational design of nanocrystal synthesis but also establish a generalizable paradigm for deciphering chemical reaction mechanisms through white-box machine learning.

Introduction

Deep learning predicts nanocrystal properties but obscures the mechanisms behind those predictions. NanoEQL addresses this limitation with a fully white-box architecture that models nanocrystal size using interpretable synthesis-related equations and scalars.

  • Motivation: Deep learning models predict nanocrystal size and optical spectra, but their black-box nature limits insight into synthesis mechanisms.The prediction process cannot be elucidated because of deep learning’s black-box nature.
  • Approach: NanoEQL is a fully white-box neural network designed to unravel nanocrystal size determination mechanisms.It builds on the EQL framework and introduces eight operators for fitting mathematical equations in nanocrystal synthesis.
  • Approach: Three smoothed operators address gradient explosion at zero for singular operators.The operators are inspired by the Michaelis-Menten equation.
  • Approach: Temperature-gated attention pooling encodes concentration-driven and reactivity-driven mechanisms while eliminating precursor sequence dependency.The strategy computes weights for different precursors and simulates the chemical reaction process.

Results and Discussion · EQL Architecture Based on Gated Attention

NanoEQL uses white-box EQL networks and temperature-gated attention to transform nanocrystal, precursor, and reaction-condition features into three interpretable scalars and predict nanocrystal size. Smoothed operators address unstable gradients near zero, helping the architecture better represent chemical reaction formulas.

  • EQL Architecture Based on Gated Attention: NanoEQL independently maps reaction conditions and nanocrystal-product features to scalars Zops and Zp, while precursor features are reduced to latent representations Zinorg and Zorg.Each feature set is processed by an independent EQL network, with inorganic and organic precursor representations reduced to four dimensions.
  • EQL Architecture Based on Gated Attention: Temperature-gated attention assigns precursor weights using concentration fractions and reactivity factors, adjusting whether pooling emphasizes concentration or reactivity.The weighted precursor representations are combined through multi-channel pooling to simulate reactions among reactants.
  • EQL Architecture Based on Gated Attention: The pooled precursor features are concatenated and processed by the grea network to derive the growth-capability scalar Zrea.The resulting scalar is combined with Zp and Zops in the model’s top-level predictor.
  • EQL Architecture Based on Gated Attention: A linear top-level predictor combines Zp, Zrea, and Zops to calculate nanocrystal size.The architecture therefore separates product, precursor, and reaction-condition processing before producing the final size output.
  • EQL Architecture Based on Gated Attention: Each EQL network uses eight operators, including identity, squaring, and reciprocal functions, followed by three hidden layers that output a scalar.The operators are intended to suffice for describing fundamental chemical reaction formulae.
  • EQL Architecture Based on Gated Attention: Standard reciprocal, square-root, and cube-root operators can update gradients excessively fast near zero, causing overuse and deviation from physical laws.This instability is identified as a limitation of the unsmoothed operator design.
  • EQL Architecture Based on Gated Attention: Three smoothed operators, inspired by the Michaelis-Menten equation, were introduced to address the near-zero gradient problem.The smoothed modifications target the reciprocal, square-root, and cube-root operators.

Performance of NanoEQL

NanoEQL predicts nanocrystal size competitively with black-box models while outperforming other genetic programming-based white-box models. Its performance improves with a linear top-level predictor and smoothed reciprocal operator, and extends to organic reaction yield prediction.

  • Nanocrystal size prediction: NanoEQL and random forest yield highly accurate predictions below 25 nm, while larger sizes show greater deviations attributed to measurement uncertainty.The deviation increases with size rather than indicating a general failure across the test range.
  • Architecture: The linear top-level predictor performs best, whereas MLP and EQL perform worst (R2=0.6), indicating that lower-layer features are effectively decoupled.A simple linear top layer also facilitates backpropagation and weight updates in lower layers.
  • Operator design: R2 increases from 0.54 to 0.65 when the smoothed reciprocal operator replaces the original reciprocal operator, preventing function abuse near prevalent zero-valued features.Zero-valued features cause abnormally large gradients near zero for the original reciprocal operator.
  • Organic reaction yield prediction: For Buchwald-Hartwig cross-coupling yield prediction, NanoEQL achieves R2 of 0.93 and MAE of 4.89%, exceeding random forest and XGBoost and approaching LightGBM (R2=0.94).This demonstrates applicability beyond nanocrystal synthesis to organic chemical reactions.

Interpretability of NanoEQL

NanoEQL represents nanocrystal size through a simple linear relationship among Zp, Zops, and Zrea. Its sparse equations reveal feature priorities spanning reaction, product, and precursor properties while exposing a performance–complexity tradeoff controlled by pruning.

  • Scalar relationship: Nanocrystal size follows a simple linear relationship with the three scalars Zp, Zops, and Zrea.These scalars provide the equation-based representation of size determination.
  • Feature importance: Reaction features are more important than product features, which outweigh reaction condition features.The relative coefficient magnitudes establish this importance hierarchy.
  • Chemical feature interpretation: The pruned networks identify distinct chemical feature groups: periodic-table positions for products, intrinsic physicochemical properties for inorganic precursors, and surface electron distribution and polarizability for organic precursors.The product network highlights rare-earth and heavy-hole elements, while precursor networks emphasize the listed physicochemical and electronic properties.

Physical Significance of the Three Scalars and Synthesis Mechanisms

NanoEQL assigns physical meaning to three scalars: -Zp measures nanocrystallization capability, Zrea measures growth capability, and -Zops varies with reaction conditions. Temperature-gated reactivity distinguishes concentration- and reactivity-driven precursor behavior and explains size changes in PbSe synthesis.

  • Three scalar meanings: Higher reaction temperatures and longer reaction times produce larger nanocrystals, while injection temperature slightly affects the -Zops potential-surface distribution.These trends connect external reaction conditions with the external input potential represented by -Zops.
  • Temperature-gated synthesis mechanisms: Inorganic precursors are predominantly concentration-driven, with an average trained reactivity coefficient of 0.18 and reactivity-driven behavior only above 300 ℃.Organic precursors have a mean reactivity coefficient of 0.33 and depend more strongly on reactivity, with concentration-driven behavior only at very short reaction times.
  • Temperature-gated synthesis mechanisms: Ligand and solvent reactivities respond differently to temperature and time: oleylamine changes substantially, oleic acid reacts mainly at high temperatures, trioctylphosphine during prolonged reactions, and octadecene only slightly.These patterns demonstrate how the temperature gate captures precursor-specific synthesis mechanisms.
  • Three scalar meanings: Zrea represents nanocrystal growth capability because it is proportional to nanocrystal size.In PbSe synthesis, increasing Pb content increases Zrea and size, whereas increasing Se content decreases both.
  • PbSe reaction-space validation: PbSe nanocrystal size decreases from 9.5 nm at a Pb:Se ratio of 1:0.25 to 6.1 nm at 1:4, consistent with Se-driven decreases in Zrea.All other reaction conditions were kept constant in the verification synthesis.

Conclusions

NanoEQL is a fully white-box neural network that combines interpretable operators and temperature-gated precursor weighting to model nanocrystal size. Its predictions support a three-scalar mechanism interpretation, practical nanocrystallization-capability screening, and transfer to organic reaction systems.

  • Conclusions: NanoEQL uses three smoothed operators to prevent singular-operator gradient explosion, preserving physical interpretability and improving predictive accuracy.The operators prevent the network from exploiting pathological gradients.
  • Conclusions: Temperature-gated attention pooling encodes concentration-driven and reactivity-driven synthesis mechanisms into precursor weighting while eliminating sequence dependency.The pooling operation simulates the chemical reaction process.
  • Conclusions: MAPE 0.35 and R2 0.65 show that NanoEQL outperforms other white-box models and matches black-box tree-based models.On the Buchwald-Hartwig reaction dataset, NanoEQL achieves R2 0.93 for yield prediction, demonstrating framework generality.
  • Conclusions: Nanocrystal size is described by a linear equation composed of -Zp, Zrea, and -Zops, representing nanocrystallization capability, growth capability, and external input potential.The three scalars can be expressed through equations derived from the original input features.
  • Conclusions: Because -Zp is independent of reactants and reaction conditions, it evaluates a material’s nanocrystallization capability and supports screening quantum dots and metal nanocrystals.Its practical utility was demonstrated through examples screening potential materials.
  • Conclusions: Average reactivity coefficients are 0.18 for inorganic precursors and 0.33 for organic precursors, indicating stronger concentration dependence for inorganic reactions and stronger reactivity dependence for organic reactions.These coefficients were obtained from analysis of the temperature-gating parameters.

Methods · Dataset and Feature Construction. · Model Architecture and Training

The study combines curated synthesis datasets and chemically informed descriptors with NanoEQL’s interpretable operator, pooling, gating, and training framework. The architecture integrates product, precursor, and reaction information to learn symbolic nanocrystal-size relationships.

  • Dataset and Feature Construction.: The nanocrystal dataset contains solution-phase synthesis recipes with resulting product size and shape, while the Buchwald-Hartwig dataset uses literature reactions and yield targets.Invalid descriptors were removed, and 15% of the data was reserved for stratified testing.
  • Dataset and Feature Construction.: Product nanocrystals and inorganic precursors were featurized with matminer, organic precursors with RDKit, and descriptors were screened using VIF<10.Nanocrystal shape was represented by circularity, aspect ratio, and vertices, then concatenated with product features.
  • Model Architecture and Training: Model inputs comprise product features, inorganic and organic precursor descriptors with molar amounts, and reaction-operation descriptors including Tinj, Trea, and t.For organic reactions, operation features instead include reaction temperature, reaction time, inert-atmosphere status, and closed-system status.
  • Model Architecture and Training: Each EQL layer expands inputs with predefined identity, polynomial, smoothed-root, reciprocal, exponential, and logarithmic operators before linear recombination.The operator library includes identity mapping, squaring, smoothed square root, smoothed cube root, smoothed reciprocal, exponential, natural logarithm, and common logarithm.
  • Model Architecture and Training: Independent early EQL subnetworks reduce precursor features to embeddings, which an attention-based pooling module aggregates using process descriptors and molar amounts.The resulting hybrid weights combine reactivity-based attention scores with normalized mass-based weights.
  • Model Architecture and Training: A temperature-dependent softmax gate uses Tinj, Trea, and t to balance mass and reactivity contributions, producing 7 pooling statistics for each precursor group.Statistics include weighted sum, weighted mean, range, weighted standard deviation, square root of weighted sum of squares, maximum, and minimum; pooled outputs generate Zrea.
  • Model Architecture and Training: The optimization objective combines Smooth L1 regression with an L1 penalty on EQL-layer weights and biases, encouraging symbolic sparsity without directly regularizing the attention scorer.Parameters are optimized with AdamW, while separate gate learning-rate treatment, Optuna tuning, warm restarts, and gradient clipping stabilize training.

Screening of Materials for -Zp Calculation

Materials Project entries were screened using composition, electronic-structure, unit-cell, and stability criteria to assemble semiconductor and metallic material sets for -Zp calculation. The screening yielded 6,800 semiconductors and 34,000 metallic materials.

  • Semiconductor screening: Semiconductor materials were retrieved from the Materials Project for screening.
  • Semiconductor screening: The semiconductor criteria excluded actinides and required bandgaps of 0.3–3 eV, direct gaps, nonmetallicity, fewer than four constituent elements, and fewer than 50 unit-cell sites.
  • Semiconductor screening: 6,800 semiconductor materials passed the screening.
  • Metallic screening: Metallic materials were screened using the actinide-exclusion, composition, unit-cell, and convex-hull criteria, yielding 34,000 materials.

Construction of the PbSe Nanocrystal Reaction Space · Synthesis and Characterization of PbSe Nanocrystals

The study uniformly generated a 56-dimensional PbSe nanocrystal reaction space by defining ranges for reaction variables, then visualized it with UMAP. For characterization, controlled hot-injection syntheses varied selenium loading while fixing other reagents and conditions, and particle sizes were measured by TEM without prior size selection.

  • Construction of the PbSe Nanocrystal Reaction Space: Reaction recipes were uniformly generated by establishing boundaries for each independent input variable.Tinj was set strictly equal to Trea.
  • Construction of the PbSe Nanocrystal Reaction Space: The reaction temperature ranged from 140 to 300 ℃, while reaction time ranged from 1 to 120 min.
  • Construction of the PbSe Nanocrystal Reaction Space: The 56-dimensional feature vector concatenated seven pooling statistics from 4-dimensional inorganic- and organic-precursor embeddings.The embeddings were denoted Zinorg and Zorg, respectively.
  • Synthesis and Characterization of PbSe Nanocrystals: All nanocrystals underwent TEM characterization without prior size selection, and average size was obtained by counting hundreds of particles.TEM observations used a FEI Tools F200S field-emission transmission electron microscope operated at 200 kV and 120 kV, respectively.

Data availability … Notes

The paper provides dataset and training-code access, supplementary analyses and figures, author contributions and affiliations, correspondence information, and a declaration of no competing interests.

  • Data availability: The dataset and NanoEQL training code are available online, while all data are provided in the main text or Supporting Information.The repository is identified as https://github.com/ime1452/Nanocrystal-.
  • Supporting Information: Supporting Information compares the eight operators, top-level predictors, reciprocal operators, and model performance for yield prediction.These comparisons appear in Tables S1–S4.
  • Supporting Information: Supporting Information reports feature explanations, dataset feature engineering, gate-factor evolution, hidden-layer weight distributions, and training/test performance across pruning ratios.These materials are provided in Tables S5 and Figures S1–S4.
  • Supporting Information: Additional analyses include -Zp sorting of 34,000 metal materials, representative examples, and potential-energy-surface distributions for -Zops.These analyses are presented in Figures S5 and S6.
  • Supporting Information: Supporting Information tracks the reactivity of oleic acid, trioctylphosphine, and octadecene with reaction temperature and time and selects training epochs using 5-fold cross-validation.The reactivity analysis is in Figure S7, and epoch selection is in Figure S8.
  • Author information: The project was conceived by K. G. and H. Z.; K. G. synthesized and characterized the nanocrystals and performed model training and evaluation.The affiliations are the MIIT Key Laboratory for Low-Dimensional Quantum Structure and Devices, School of Materials Sciences & Engineering, Beijing Institute of Technology, Beijing.
  • Author information: K. G. and H. Z. analyzed the models and wrote the manuscript; correspondence is directed to Kai Gu or Haizheng Zhong, and the authors declare no competing interests.These statements appear in the author-information and notes sections.
Loading 2608.14734v1…