Source-linked AI summary
Prototype-guided transfer of sparse literature knowledge for electrolyte additive discovery
Weixiang Hong, Hongting Du, Jiayue Tang, Ruifeng Tan, Yangjian Quan, Jia Li, Jiaqiang Huang
TL;DR
Electrolyte additive discovery must prioritize candidates despite sparse reported successes and vast, largely unlabeled chemical spaces. ProtoMI learns transferable prototypes from reported additives and adapts them to candidate space, enabling experimentally efficient prioritization. TNDB improved high-temperature LiFePO4||graphite cycling relative to the baseline electrolyte while analyses indicated modified inorganic interphases and reduced solvent decomposition.
Problem
Successful additives are sparsely reported, while most candidate molecules occupy a much larger unlabeled chemical space, making candidate prioritization difficult.
Method
ProtoMI converts sparse literature-derived additive knowledge into transferable molecular prototypes and adapts those prototypes to unlabeled chemical space for candidate exploration.
Results
34.93% improvement in high-temperature LiFePO4||graphite cycling relative to the baseline electrolyte was obtained with representative candidate TNDB.
Takeaways & Limitations
ProtoMI recommendations can be connected to battery performance, with TNDB associated with B-containing and F/P/O-modified inorganic interphases, suppressed carbonate-solvent decomposition, and reduced Fe deposition on graphite.
Takeaways & Limitations
The current implementation focuses on boron-containing additives and does not yet explicitly incorporate electrochemical stability and reaction behavior into its representations.
Abstract
from arXiv · showhide
Electrolyte additive discovery remains challenging because experimentally validated molecules are sparse, whereas accessible chemical spaces are vast and largely unlabeled. This challenge is amplified in lithium-ion batteries, where additive performance arises from coupled interfacial reactions rather than a single molecular property. Here, we develop a prototype-guided molecular intelligence, ProtoMI, a literature-driven framework that learns transferable structural priors from reported electrolyte additives and uses them to prioritize candidates in unlabeled chemical space. For boron-containing additives, ProtoMI combines 126 literature-reported molecules with 179,977 unlabeled candidates. Graph contrastive learning identifies seven chemically interpretable prototypes from the reported additives, and prototype guided semi-supervised contrastive learning adapts these prototypes to the candidate space under source-target distribution mismatch. In retrospective temporal validation, ProtoMI achieves enrichment factors of 9.2-45.6 while screening less than 2% of the candidate space. A subsequent translation step identifies four commercially accessible candidates. One representative candidate, 4,4,5,5-Tetramethyl-2-[10-(1naphthyl)anthracen-9-yl]-1,3,2-dioxaborolane (TNDB), improves high-temperature LiFePO4||graphite cycling at 55 °C by 34.93% relative to the baseline electrolyte. An arsenal of characterizations and operando optical fiber Fourier transform infrared spectroscopy suggest that TNDB forms B-containing, F/P/O-modified inorganic interphases, suppresses solvent decomposition and reduces Fe deposition on graphite. This case study shows how sparse literature knowledge can guide experimentally efficient molecular discovery in data-scarce battery-additive spaces.
Results
ProtoMI extracts chemically interpretable prototype structure from sparse reported additives and adapts it to broader unlabeled boron-containing chemical space. The adapted prototypes preserve meaningful structural organization while supporting coherent, diverse candidate prioritization.
- Dataset construction: 94% accuracy for additive identification was achieved by the LLM-assisted extraction pipeline, which required approximately 11 s per article and less than 60 CNY total processing cost.
- Dataset construction: Reported additives occupied restricted, recurring regions of a broader unlabeled chemical space rather than being uniformly distributed.The positive and unlabeled sets also differed in molecular size, connectivity, and local chemical environments.
- Prototype discovery: Graph-based encoders achieved silhouette scores of 0.751-0.832, exceeding fingerprint- and SMILES-based representations with scores of 0.258-0.337.The GINE+ model achieved the strongest organization, while augmentation-enhanced representations further improved embedding structure.
- Prototype discovery: Seven molecular prototypes captured distinct combinations of recurring boron-related fragments and chemically meaningful structural preferences.Examples included oxalate and tetra-alkoxy frameworks in P3, fluorinated aromatic fragments in P2, and catechol-derived structures in P4.
- Prototype adaptation: Adapted prototypes reorganized literature-derived knowledge according to the target candidate space while retaining structural continuity and characteristic boron-containing motifs.Prototype drift decreased through exploration, refinement, and stabilization, indicating convergence toward a stable organization used for candidate prioritization.
- Recommendation performance: 0.57 entropy-weighted prototype similarity was achieved while recommended molecules remained structurally coherent and distributed across multiple prototype regions.The framework reported prototype similarity of 0.789 ± 0.030, assignment entropy of 0.723, and cluster separability of 0.976, exceeding competing recommendation pipelines on entropy-weighted similarity.
- Ablation analysis: Removing top-k assignment, prototype decorrelation, or EMA updating degraded confidence, prototype separation, or cross-domain adaptation stability.
M9. Hierarchical clustering algorithm with outlier filtering. To identify
The study clusters learned molecular graph embeddings to construct prototypes, filters outliers, and evaluates prototype separation before downstream transfer and electrochemical validation. It also defines contrastive objectives and describes the cell-testing and characterization protocols.
- Clustering and filtering: Hierarchical clustering uses learned molecular graph embeddings with average linkage and cosine distance to construct molecular prototypes.The resulting filtered embeddings, labels, additive names, and linkage matrix are retained for prototype construction and prototype-guided representation learning.
- Clustering and filtering: Outliers are removed using cluster-specific cosine-distance thresholds set at the mean distance plus 1.5 standard deviations.Clusters containing fewer than two samples are excluded from this filtering procedure before clustering is recomputed.
- Cluster selection and evaluation: The silhouette score with cosine distance is evaluated across candidate cluster numbers, and the highest-scoring configuration is selected.Evaluation uses prototype-relevant core samples selected from held-out positive and unlabeled test molecules by the same top-k criterion.
- Contrastive learning: InfoNCE contrastive learning encourages augmented views of each molecular graph to be more similar than other batch samples treated as negatives.Embeddings are L2-normalized and compared using cosine similarity, with temperature controlling distribution concentration.
- Contrastive learning: Prototype semi-supervised contrastive learning treats similarity to the assigned prototype as a positive logit and similarities to other prototypes as competing logits.The total objective combines prototype contrastive loss with prototype decorrelation loss to encourage diverse molecular patterns.
Competing interests
The authors report a pending Chinese patent application related to the work and provide source-code access, while datasets are available from the corresponding author upon reasonable request.
- Competing interests: A Chinese patent application related to the work has been filed and is currently pending.The source code is publicly available, while generated and analyzed datasets are available from the corresponding author upon reasonable request.
Supplementary information
The supplementary information documents the reported-additive dataset, clustering and prototype analyses, molecular motifs, and prototype-transfer behavior. It also provides implementation details for molecular features and similarity-based recommendation evaluation.
- Reported-additive organization: Hierarchical clustering of 126 reported boron-containing additives reveals recurrent structural families including lithium borates, fluorinated borates, and aryl borate derivatives.The organization supports recurring scaffold-level motifs rather than isolated successful examples.
- Supplementary tables: The supplementary materials include dataset statistics, node-feature meanings, reported additives by prototype, shared motifs, and prototype summaries.These resources are listed in Tables S1–S5.
- Prototype discovery: Silhouette analysis identifies seven as the optimal number of molecular prototypes, with hierarchical clustering achieving its highest score at k = 7.K-means shows a consistent optimum in the same region.
- Prototype discovery: The learned representation organizes reported additives into seven structurally and functionally distinguishable prototype regions.UMAP visualizes graph-level embeddings colored by prototype assignment, and the spatial separation is consistent with silhouette-score analysis.
- Prototype transfer: The prototype-correspondence heatmap shows strong diagonal similarity between initial and transferred centroids, alongside structured off-diagonal similarities.This indicates preserved prototype identities with partial reorganization and information sharing during adaptation.
- Prototype transfer: Transferred prototypes retain chemically meaningful motif enrichment after adaptation rather than forming arbitrary candidate-space clusters.Analyzed motifs include tetra-alkoxy borate, tri-alkoxy borate, boron oxalate, BF3-like motifs, fluorinated alkoxy groups, C-B(OR)2, pentafluorophenyl, and catechol boronate.
- Recommendation evaluation: ProtoMI recommendations show the highest overall molecule-level prototype similarity among the compared translation strategies.The distribution is summarized with violin plots, individual points, and median/interquartile-range box plots.