Source-linked AI summary
When do machine-learned exchange-correlation improvements inherit into density-functional tight binding?
Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban
TL;DR
The paper asks how much improvement from machine-learned exchange-correlation functionals survives density-functional tight-binding parameterization despite a structural mismatch between orbital-dependent operators and multiplicative potentials. Using a transfer-ratio pre-test and targeted controls, it finds that band-gap corrections anti-transfer, while occupied-manifold properties, repulsive potentials, and some oxide gaps inherit selectively.
Problem
How much of a parent functional’s improvement survives tight-binding parameterization remains unclear despite available code paths, because orbital-dependent functionals exceed the channel’s multiplicative-potential representation.
Method
The study evaluates a transfer ratio on small test solids and uses practical inversion and functional controls to assess which parent-level properties the parameterization channel can carry.
Results
Band-gap corrections anti-transfer, while occupied-manifold properties, ionic and closed-shell repulsive potentials, and rocksalt-oxide gaps inherit selectively across material classes.
Takeaways & Limitations
The transfer ratio offers a cheap pre-test for deciding whether a parent functional is suitable for a tight-binding parameterization campaign.
Takeaways & Limitations
The representational obstruction constrains the functional class, so exchanging one machine-learned functional for another does not remove it.
Abstract
from arXiv · showhide
Machine-learned exchange-correlation functionals correct band gaps at near-semilocal cost, while density-functional tight binding reaches the $10^3$-$10^6$-atom regime; combining them assumes that a better parent yields a better parameterization, but we show it does not. Current-generation functionals are orbital-dependent generalized Kohn-Sham operators, whereas the parameterization channel is built on a multiplicative potential, preventing exact representation. Using the transfer ratio, the surviving fraction of a parent-level change, we find anti-transfer: coherently negative ratios across four covalent semiconductors move the gap in the wrong direction, consistent with a molecular proxy and an r$^2$SCAN control. The minimal-basis overgap is dominated by the on-site convention rather than basis incompleteness; correcting the on-site block removes most of it, while one $d$-polarization shell closes a further $16$-$40%$, depending on the placement of the empty $d$ level, which no free-atom eigenvalue uniquely fixes. Occupied-manifold enhancements, ionic and closed-shell repulsive potentials, and rocksalt-oxide gaps inherit, whereas elemental and III-V covalent networks inherit neither gaps nor repulsive potentials and oxide networks inherit only the latter. We screen 23 elements and release the parameter sets, showing that the transfer ratio provides a cheap pre-test before any parameterization campaign.
1 Introduction
DFTB compresses DFT to reach systems far beyond routine DFT scale, motivating tests of whether machine-learned functional improvements survive parameterization. This study defines a transfer ratio and shows that orbital-dependent functionals cannot be represented exactly through the standard multiplicative-potential channel.
- Motivation: DFT is limited to a few hundred atoms, whereas DFTB reaches the 10^3-10^6-atom regime at roughly three orders of magnitude lower cost.DFTB achieves this through a minimal valence basis, two-center Slater–Koster tables, and a local effective potential.
- Motivation: The central untested assumption is that a more accurate parent functional yields a better DFTB parameterization.The assumption is important because machine-learned tight-binding work increasingly transfers conclusions from semilocal calculations to higher rungs.
- Representation barrier: Orbital-dependent machine-learned functionals are generalized Kohn–Sham operators, but standard DFTB parameterization consumes a multiplicative local potential.Therefore, the current parameterization channel cannot represent these functionals exactly.
- Study design: The transfer ratio measures the fraction of a parent-DFT property change that survives in DFTB when only the functional changes.A potential-inversion bridge admits any functional into a confined-atom pipeline under one overlap-certified convention.
- Study design: Most unfitted minimal-basis overgap arises from the on-site convention rather than basis deficiency, while one added d-polarization shell closes part of the remainder.The remaining closure depends on the placement of the empty d level, and angular-channel inconsistency cannot be repaired within a shared local potential.
2 Results
Machine-learned parent-functional gap corrections do not transfer through the density-functional tight-binding parameterization channel: across covalent semiconductors they anti-transfer because orbital-dependent effects lack an exact multiplicative-potential representation. The transfer ratio exposes this failure, while correcting the on-site convention removes most minimal-basis overgap inflation.
- Transfer ratio: The transfer ratio measures the tight-binding change divided by the parent-DFT change, with 1 denoting full inheritance and 0 denoting none.It is intended as a pre-test before committing to a parameterization campaign.
- Band-gap transfer: Under the free-atom on-site convention, gap corrections anti-transfer coherently across diamond, silicon, 3C-SiC, and boron phosphide, with mean −0.97 and range 0.31.The useful-transfer decision threshold is 0.5; the parent functional opens the gap, while the tight-binding response is small and has the wrong sign.
- Occupied-manifold transfer: The occupied bandwidth also fails the 0.5 transfer gate, averaging +0.009 with a range of 0.963 across the three reporting solids.It fails less severely than the gap but much less coherently, while occupied-manifold widths are theoretically among the effects that can transfer.
- Channel representability: Orbital-dependent functional effects acting through orbital curvature cannot be represented exactly by the multiplicative-potential channel, whereas density-mediated potential effects can transfer.The discarded contribution is a gradient overlap weighted by f, and exchanging one machine-learned functional for another does not remove this structural limitation.
- Basis and on-site convention: Correcting the on-site convention collapses the silicon Γ-point overgap by 9.5 to 15.4 eV, showing that most minimal-basis inflation is bookkeeping rather than basis incompleteness.The minimal single-s–p parameterization gives 11.81 eV versus a parent direct gap of 2.28 eV; corrected values bring diamond and 3C-SiC close to experiment but overcorrect silicon.
3 Discussion
The standard parameterization channel preserves occupied-manifold and orbital-geometry information but discards fundamental-gap corrections driven by orbital dependence. The transfer ratio therefore serves as a low-cost pre-test, while basis and Hamiltonian extensions offer partial routes toward improved transfer.
- Transferability: The transfer ratio tests, at negligible cost, whether a parent-functional property is likely to survive tight-binding compression.Compression preserves occupied-manifold and orbital-geometry information but discards information residing in the fundamental gap.
- Representational barrier: Orbital dependence exceeds the standard channel’s representational capacity, making gap improvement and compressibility fundamentally opposed when it carries most of the correction.An optimized effective potential provides the best single multiplicative representative but cannot restore the orbital-dependent response.
- Representational barrier: A conventional meta-GGA anti-transfers in every solid under both on-site conventions, showing that the obstruction is associated with orbital dependence rather than machine learning.The r2SCAN control establishes that machine-learned parents are not uniquely responsible for anti-transfer.
- Basis and on-site effects: One to two electronvolts of residual overgap remains after on-site bookkeeping is corrected, and a polarization shell closes only part of it because the empty d-level placement controls the result.This separates a largely repairable on-site convention error from a residual basis limitation whose correction is partial and placement-dependent.
- Chemical scope: Ionic and closed-shell repulsive potentials transfer cleanly, rocksalt-oxide gaps approach their target, and covalent networks require Hamiltonian or basis extensions for comparable gap transfer.Oxide networks already inherit their repulsive potentials, whereas the strongest footing for functional-consistent tight binding is closed-shell oxide chemistry.
- Scope and future directions: These conclusions apply to unfitted, functional-consistent parameterizations, not directly to data-driven tables fitted to reference bands or to empirically tuned DFTB.The benchmark covers four covalent solids and one ionic control, with a molecular proxy of three diatomics, so reported averages carry limited scope.
4 Methods
The study combines confined-atom Kohn–Sham inversion, periodic and molecular parent calculations, and DFTB parameterization to measure functional-change transfer under controlled settings. Deterministic convergence tests, held-out repulsive-potential validation, and released parameter sets support reproducibility.
- Atomic calculations: Confined pseudo-atoms were solved with a PySCF radial Kohn–Sham solver using level-6 grids, even-tempered valence bases, confinement potentials, and scalar-relativistic treatment beyond sulfur.Confinement used r0 = 1.85 rcov, with hydrogen fixed at 3.0 Bohr and silicon at 3.30 Bohr; silicon included a 3d polarization channel.
- Potential construction: Functional effective potentials were assembled directly when possible or recovered by radial Kohn–Sham inversion with node preservation, spherical-harmonic projection, and a 5% orbital-amplitude node mask.The mask retained radii where both |us| and |up| exceeded 5% of their maxima, excluding ill-conditioned regions.
- DFTB calculations: DFTB calculations used DFTB+ with solver variants for molecular, large-cell, and ΔSCF calculations, while self-consistent charges were converged to 10^-5 e on the Mulliken charge.Dense ELPA diagonalization was assigned to defect quantities and the linear-scaling NTPoly solver to gapped bulk.
- Transfer measurements: The transfer ratio was defined as the DFTB property change divided by the parent-DFT property change and evaluated for gaps and occupied bandwidths across covalent, ionic, and molecular systems.Periodic calculations used fixed Materials Project lattice constants, including a = 3.567 Å for diamond, 5.430 Å for silicon, and 4.360 Å for 3C-SiC.
- Repulsive potentials: Repulsive potentials were fitted with curvature-constrained splines to rattled cells, equation-of-state scans, and dimer bond scans, then scored on held-out configurations using centered RMS error and Pearson correlation.The train/test split and mean-centering were used to remove offset dependence from transferability assessment.
- Reproducibility and resources: All calculations were deterministic, with numerical and systematic uncertainty bounded through basis/grid convergence and a carbon confinement-radius sweep from 2.0 to 7.0 Bohr.The project ledger totaled 290 GPU jobs and 92.5 active GPU-hours; released data, parameter sets, manifests, and validation workflows are openly available.
Supplementary Information
This section examines when machine-learned exchange–correlation improvements transfer into density-functional tight binding.
- The section concerns machine-learned exchange–correlation improvements.
- The section addresses their inheritance into density-functional tight binding.
Supplementary Note 1 Derivations · Supplementary Note 1.1 Notation · Supplementary Note 1.2 Assumptions
The supplementary note provides complete derivations for the main-text theoretical results and states the assumptions underlying them. It emphasizes that the parameterization-channel error is an equality, while identifying idealizations and artifacts that require direct validation.
- Supplementary Note 1 Derivations: The note supplies complete derivations supporting the main-text theoretical results, with the parameterization-channel error given as an equality rather than a bound.This equality makes the resulting predictions sharp.
- Supplementary Note 1.2 Assumptions: Every theoretical result rests on explicitly stated assumptions presented before the proofs.The assumptions define the conditions under which the derivations apply.
- Supplementary Note 1.2 Assumptions: The confined atom is assumed spherical and spin-unpolarized, with radial potentials and orbitals separated into R_nl(r)Y_lm.This is assumption A1.
- Supplementary Note 1.2 Assumptions: The exchange-correlation energy density is differentiable in τ, f = ∂e_xc/∂τ is continuous, and R_nl are exact radial generalized Kohn–Sham solutions.These are part of assumptions A2 and A3.
- Supplementary Note 1.2 Assumptions: The boundary condition fϕ_μ∇ϕ_ν → 0 removes surface terms in integration by parts for confined orbitals vanishing outside the confinement radius.This is assumption A4.
- Supplementary Note 1.2 Assumptions: Assumption A3 is the main idealization because differing radial nodes make |v_s−v_p| combine physical l-dependence with numerical artifacts.The s channel has a radial node whereas the p channel does not.
- Supplementary Note 1.2 Assumptions: Prediction P1 is required because under PBE the physical term vanishes identically, making residual |v_s−v_p| a direct artifact measure.The note also states that A2 fails at a nucleus for some functionals.
Supplementary Note 1.3 The generalized Kohn–Sham equation · Supplementary Note 1.4 Proof of Proposition 3
The supplement derives the generalized Kohn–Sham equation for orbital-dependent functionals and proves its reduction to an l-resolved local representative under stated assumptions. Uniqueness follows because the ordinary radial equation fixes that multiplicative potential pointwise once the orbital and eigenvalue are fixed.
- Supplementary Note 1.3 The generalized Kohn–Sham equation: Collecting multiplicative contributions yields the generalized Kohn–Sham equation used in the main text.
- Supplementary Note 1.3 The generalized Kohn–Sham equation: The resulting equation is the standard orbital-dependent generalized Kohn–Sham form.
- Supplementary Note 1.3 The generalized Kohn–Sham equation: It is equivalently a position-dependent effective-mass problem with m*(r) = [1 + f(r)]^-1.
- Supplementary Note 1.4 Proof of Proposition 3: Under A1, separating ψ = R_nl(r)Y_lm and setting g = 1 + f reduces the analysis to a radial equation for radial g.
- Supplementary Note 1.4 Proof of Proposition 3: For radial g, the generalized Kohn–Sham equation reduces to the radial form used in the proof.
- Supplementary Note 1.4 Proof of Proposition 3: The parameterization channel expresses the ordinary radial Kohn–Sham equation, corresponding to Eq. (9) with g ≡ 1 and a multiplicative v_l.
- Supplementary Note 1.4 Proof of Proposition 3: Subtracting the ordinary equation from the generalized one and dividing by R_nl away from its nodes produces the l-resolved local representative.
- Supplementary Note 1.4 Proof of Proposition 3: Uniqueness is immediate because the ordinary radial equation determines v_l pointwise when R_nl and ε_nl are fixed.
Supplementary Note 1.5 Affine structure of the representative · Supplementary Note 1.6 Proof of the non-representability theorem
The representative is exactly affine in angular momentum at fixed radial factor, because all quadratic l-weighting cancels. The non-representability proof shows that a common channel representative would require a scalar identity not satisfied by current functionals or universally across elements.
- Supplementary Note 1.5 Affine structure of the representative: At fixed S_nl, v_l − v is exactly affine in l, with C depending on S_nl but not l.This follows by substituting the cusp-regular form into Eq. (3) and differentiating with respect to l.
- Supplementary Note 1.5 Affine structure of the representative: The local representative’s l(l + 1) terms cancel identically, leaving no quadratic l-weighting; surviving l-dependent terms are proportional to f or f′.The cancellation occurs between the explicit kinetic-like term and the corresponding contribution generated by R′_nl.
- Supplementary Note 1.5 Affine structure of the representative: The affine structure predicts a factor of two between v_d − v_s and v_p − v_s, rather than the factor of three implied by a naive l(l + 1)-weighted reading.The discrepancy arises because the apparent quadratic weighting is canceled exactly.
- Supplementary Note 1.6 Proof of the non-representability theorem: If f ≡ 0, the local representative gives v_l = v for every l, so v_ch = v satisfies Definition 2.This covers the local and generalized-gradient rungs underlying the parameterization channel for thirty years.
- Supplementary Note 1.6 Proof of the non-representability theorem: The cusp expansion fixes the near-nucleus radial behavior as R_nl(r) = r^l(1 − Zr/(l + 1) + O(r^2)) up to normalization.This provides the leading-order condition used in the non-representability argument.
- Supplementary Note 1.6 Proof of the non-representability theorem: The corrected near-nucleus law vanishes for l = 0 but generally remains nonzero for l ≥ 1 unless its bracket vanishes.The standard condition is recovered only when f(0) = f′(0) = 0; measured carbon corrections shift coefficients by 0.46% at l = 1 and 0.72% at l = 2.
- Supplementary Note 1.6 Proof of the non-representability theorem: A common channel representative would force f′(0) = Zf(0), a scalar identity not satisfied by any functional in use or by one functional for every element.Because Z varies across elements, while f(0) and f′(0) are determined by the functional’s evaluated density, the condition cannot hold universally.
Supplementary Note 1.7 The coefficients, and why three channels close the escape · Supplementary Note 1.8 Proof of the channel-residual identity
The supplementary proofs show that three angular channels eliminate any exact local multiplicative-potential representative, while the channel residual is instead a kinetic-matrix rescaling. The resulting gap-shift analysis fixes the missing physics’ sign and scale but carries explicit first-order, self-consistency, and basis limitations.
- Supplementary Note 1.7 The coefficients, and why three channels close the escape: The s-channel coefficient vanishes unconditionally, while p and d coefficients vanish at κ = 1 and κ = 2/3, respectively.Because no single κ annihilates both p and d coefficients, an atom with s, p, and d channels admits no exact local representative for any functional.
- Supplementary Note 1.7 The coefficients, and why three channels close the escape: The ratio 3(2 − 3κ)/(1 − κ) equals 4/3 only at κ = 0, so the l/(l + 1) weights are not parameter-free.The naive l(l + 1) interpretation would force the ratio to 3 for every κ and is excluded by the stated cancellation.
- Supplementary Note 1.7 The coefficients, and why three channels close the escape: The obstruction arises from diverging near-nucleus channel representatives that no bounded multiplicative potential can reconcile.The theorem relies on the exact radial cusp condition, which fails for Gaussian-basis orbitals without a cusp.
- Supplementary Note 1.8 Proof of the channel-residual identity: If a channel residual were representable by a multiplicative w for every basis pair, orbitals with identical density but different gradients would yield a contradiction unless f were constant.The residual’s dependence on gradients is therefore the core reason a local potential representation fails.
- Supplementary Note 1.8 Proof of the channel-residual identity: The residual equals 1/2 f⟨∇ϕμ|∇ϕν⟩, a kinetic-matrix rescaling rather than a potential, and cannot be absorbed by confinement, superposition, or inversion changes.This establishes the channel-residual identity’s central nonlocal obstruction.
- Supplementary Note 1.8 Proof of the channel-residual identity: The sign argument predicts greater conduction-edge than valence-edge curvature in covalent bonding regions, but it is not a proof and does not directly extend to ionic solids.The inequality is local, the integral spans the cell, and f is not sign-definite everywhere.
- Supplementary Note 1.8 Proof of the channel-residual identity: The first-order gap-shift expression fixes the sign and location of missing physics, not its exact magnitude; self-consistent charge can restore part of the shift.It is an upper bound on lost physics only in the non-self-consistent limit, and the confined orbitals also change when the functional changes.
- Supplementary Note 1.8 Proof of the channel-residual identity: 10^-2 to 10^-1 Ha is the expected s p σ residual scale, given dimensionless f of order 10^-1 and kinetic-scale gradient overlaps of order 10^-1 to 1 Ha.The apparent node-region extremum belongs to the diagnostic inversion, not the released parameterization or reported gaps and transfer ratios.
Supplementary Note 1.10 Scope of the theorem
The theorem extends to orbital-dependent channels beyond τ, but its l-dependence prediction is limited to meta-GGA parents rather than local hybrids. It concerns parent-derived tables, with residual anti-transfer and minimal-basis overgap barriers remaining even in best-case representations.
- Scope beyond τ: Orbital dependence extends the channel-representability argument beyond τ to local hybrids and density-matrix features, whose functional derivatives are nonmultiplicative.For local hybrids, the derivative includes an exchange-kernel term, while CIDER24Xe uses density-matrix features.
- Scope beyond τ: The l-dependence law applies specifically to the τ channel, so prediction P2 covers meta-generalized-gradient parents but not local hybrids.The residual’s closed form changes outside τ, while the l-dependence law does not carry over to local hybrids.
- Qualifications: The theorem constrains tables derived from a parent functional, not models fitted directly to a better functional’s bands.Channel representability does not apply to band-fitted models.
- Qualifications: Residual ΔH remains even for the best local representative, making the measured anti-transfer a best-case statement.The theorem’s representability limit does not eliminate the residual Hamiltonian error.
- Qualifications: The minimal-basis overgap is an independent, partly repairable barrier that does not follow from the theorem’s channel-representability argument.This separates the overgap limitation from the theorem’s derivational scope.
Supplementary Note 3 The bridge derivation and the ruled-out routes
The bridge cannot directly represent meta-GGA functionals because its two-center integrator uses density and gradient only, while convention-dependent on-site energies create large gaps and anti-transfer. The ruled-out routes therefore identify local-potential incompatibility and on-site conventions as decisive limitations, with polarization effects requiring a separately chosen d level.
- Bridge derivation and ruled-out routes: The bridge integrator rejects meta-GGA functionals because its libxc wrapper uses only density and gradient, not kinetic-energy density, preventing r2SCAN two-center tables.In the r2SCAN control, the functional enters only through meta-GGA-capable atomic on-site eigenvalues.
- Bridge derivation and ruled-out routes: No bound free-atom d level exists for polarization, so the extended-basis tables retain confined-atom d on-site energies: carbon +1.5054 Ha and silicon +0.9812 Ha.The confined value is the only d energy supplied by the generating atomic solve.
- Bridge derivation and ruled-out routes: Free-atom rebuilding collapses diamond and silicon CIDER23X gaps from 15.29 → 5.793 eV and 11.09 → 0.156 eV, removing 9.5–15.4 eV of overgap.This confirms the on-site convention as the dominant source of the minimal-basis overgap.
- Bridge derivation and ruled-out routes: After convention correction, free-atom transfer ratios are −0.771, −1.051, −1.085 and −0.974 across the four covalent solids, indicating anti-transfer.The full l-resolved inverted-potential comparison supports this as the primary result.
- Bridge derivation and ruled-out routes: The carbon s pσ integral changes sign between functionals, from −0.249 Ha for semilocal to +0.283 Ha for machine-learned, demonstrating failure of a single local potential.The corresponding failure is shown qualitatively for carbon and quantitatively for silicon.
Supplementary Note 4 Atomic reference, convergence, and the solver
The supplementary note establishes converged atomic references, transfer calculations, and solver settings, while documenting thousand-atom silicon performance and ionic cases where confinement reverses or exaggerates the parent gap.
- Atomic reference: The confined-atom basis was enlarged until free-atom energies stabilized below a microhartree, with heavier elements plateauing at 28 functions.The atomic Hubbard U is stored in the U_s slot with U_s = U_p, following the 3ob convention.
- Convergence: Transfer-solid gaps and transfer ratios were converged across numerical pipeline parameters, reproducing frozen production gaps to sub-0.002 eV.Sweeps covered parent plane-wave cutoffs and k-meshes, plus DFTB two-center grids and self-consistent-charge tolerances.
- The solver: ELPA, NTPoly, and QR agreed to full precision at 1.8841742541 Ha, while ELPA remained faster than NTPoly for amorphous silicon up to 1728 atoms.At 216 atoms, ELPA provided a ~22× speedup; total energies agreed to ~1 mHa.
- The solver: A 1000-atom silicon supercell converged in 6.9 s on a single CPU with a size-independent gap.The inter-functional difference was decoupled by the largest cell.
- Ionic exceptions: Sodium fluoride undergapped at 1.46 eV versus a 10.67 eV parent gap, whereas magnesium sulfide overgapped at 11.56 versus 7.09 eV.The sodium-fluoride undergap was attributed to fluorine’s tight 1.99 Bohr confinement radius, while sodium chloride inherited because chlorine’s radius was larger.
- Transfer-solid provenance: Boron phosphide and magnesium oxide reused validated CIDER23X two-center tables and followed the same parent-to-DFTB pipeline in the free-atom on-site convention.No new CIDER table generation was required for either solid.
Supplementary Note 6 Channel-potential asymmetry and the basis de-
The note quantifies channel-potential non-representability through |v_s−v_p| extracted by radial Kohn–Sham inversion. Node masking removes inversion artifacts, while PBE provides a measured 0.006–0.061 Ha noise floor; near-nucleus fits are stable for carbon, nitrogen, and oxygen.
- Channel-potential asymmetry: |v_s−v_p| quantifies theorem non-representability by inverting the confined-atom radial Kohn–Sham equation separately for each angular-momentum channel.The channel potentials are reconstructed from the radial functions and eigenvalues.
- Channel-potential asymmetry: 103–104 Ha is the unmasked pointwise |v_s−v_p| at radial nodes, identifying fixed-radius s p σ values as inversion artifacts without physical meaning.The divergence occurs where the radial function approaches zero, so percentiles are reported only on a node mask.
- Channel-potential asymmetry: 0.006–0.061 Ha is the measured PBE median |v_s−v_p|, serving as the inversion noise floor despite PBE requiring |v_s−v_p| = 0.Table S3 reports 10th, 50th, and 90th percentiles on a 500-point, 0.02–6.0 Bohr grid with a 350-direction spherical average.
- Near-nucleus fit: 1.5% is the maximum reproducibility variation of f(0) across near-nucleus windows for carbon, nitrogen, and oxygen, with spreads of 0.5%, 0.9%, and 1.2%.The fitted slope is stable across r < 0.02 and r < 0.05 Bohr windows; wider windows are excluded by curvature.
Supplementary Note 7 The extended-basis test and the two table families
Supplementary Note 7 separates authoritative released parameter sets from electronic-only extended-basis tables and tests basis effects under a controlled rebuild. Adding d polarization removes most of the minimal-sp overgap, but the closure depends on the empty-d on-site convention, while the silicon–oxygen branch exhibits sign-inverted gap transfer.
- Two table families: The released tables are authoritative for gaps, transfer ratios, and overgaps, while electronic-only tables serve only the extended-basis sp-versus-spd decomposition.The two table families agree within ≤0.01 eV for diamond and silicon, but differ for 3C-SiC, giving 6.37 and 2.22 eV versus 6.847 and 2.915 eV.
- Extended-basis test: The rebuild fixes confinement radii and on-site eigenvalues, isolating the valence l-set change; the minimal-sp pipeline reproduces the frozen baseline.Silicon’s Γ-gap is 11.811 versus 11.808 eV, while diamond is reproduced exactly at 14.973 eV.
- Empty-d convention: 25.8%, −5.5%, and 28.7% are the per-solid closures when confined-atom empty-d levels are retained under the free-atom on-site convention.The empty d channel has no bound free-atom level, so closure depends on how that channel is treated; the values correspond to carbon, silicon, and the third tested solid in the cited convention.
- Silicon–oxygen branch: −3.368 eV is the corrected silicon–oxygen gap change, sign-inverted relative to the parent after an earlier oxygen–oxygen fitting error inflated it to 3.54 eV.The silicon–oxygen Slater–Koster set was built through the same pipeline because no reference parameterization exists.
Supplementary Note 10 Compute footprint and released parameter sets
The supplementary note reports the project’s compute footprint and releases three Slater–Koster parameter sets generated under an identical automated protocol. The inversion bridge is a one-time per-element solve, while the ledger is dominated by periodic parent-DFT reference gaps.
- Compute footprint: 290 GPU jobs ran across five NVIDIA RTX 4090 nodes, spanning 92.48 h actively and 118.72 h in summed job-wall time.Summed job-wall time exceeds active span because jobs ran concurrently on one node.
- Compute footprint: The inversion bridge is solved once per element and amortized across downstream two-center tables and every crystal or molecule built from that element.Project totals are dominated by periodic parent-DFT reference gaps.
- Released parameter sets: The release contains three Slater–Koster parameter sets produced under one identical automated protocol, varying only the parent functional.Each set includes homonuclear on-site blocks and heteronuclear Hamiltonian and overlap tables.
- Released parameter sets: The released machine-learned functionals use CiderPress 0.4.0 definitions bridged into the confined-atom solver by signed-orbital Kohn–Sham inversion.The definitions are identified as Zenodo 13336814.