Source-linked AI summary

Polymer Informatics: Current Status and Critical Next Steps

Lihua Chen, Ghanshyam Pilania, Rohit Batra, Tran Doan Huan, Chiho Kim, Christopher Kuenneth, Rampi Ramprasad

arXiv:2011.00508v1cond-mat.softcs.LG

TL;DR

Polymer informatics addresses the challenge of searching polymers’ vast chemical space amid limited curated data and incomplete machine-readable representations. The paper reviews data generation, polymer representations, AI/ML prediction and design, synthesis validation, and emerging ecosystem components. Reported examples show successful candidate discovery and more efficient active learning, while polymer representations and multi-fidelity methods retain important limitations.

  • Problem

    Polymer informatics lacks widespread curated data and representations that capture polymer structure together with synthesis and processing conditions.

  • Method

    The paper reviews polymer data resources, representations, machine-learning prediction and design methods, applications, and critical next steps.

  • Results

    Machine-learning-assisted design discovered polymers with targeted application properties, including two CO2/CH4-separating membranes from over 11,000 homopolymers.

  • Takeaways & Limitations

    Integrating polymer data, representations, surrogate models, interfaces, and computational or synthesis validation can support polymer discovery and design.

  • Takeaways & Limitations

    Graph neural-network use remains limited because representing large-scale polymer morphology and repeat-unit connection points is difficult.

Abstract

from arXiv · show

Artificial intelligence (AI) based approaches are beginning to impact several domains of human life, science and technology. Polymer informatics is one such domain where AI and machine learning (ML) tools are being used in the efficient development, design and discovery of polymers. Surrogate models are trained on available polymer data for instant property prediction, allowing screening of promising polymer candidates with specific target property requirements. Questions regarding synthesizability, and potential (retro)synthesis steps to create a target polymer, are being explored using statistical means. Data-driven strategies to tackle unique challenges resulting from the extraordinary chemical and physical diversity of polymers at small and large scales are being explored. Other major hurdles for polymer informatics are the lack of widespread availability of curated and organized data, and approaches to create machine-readable representations that capture not just the structure of complex polymeric situations but also synthesis and processing conditions. Methods to solve inverse problems, wherein polymer recommendations are made using advanced AI algorithms that meet application targets, are being investigated. As various parts of the burgeoning polymer informatics ecosystem mature and become integrated, efficiency improvements, accelerated discoveries and increased productivity can result. Here, we review emergent components of this polymer informatics ecosystem and discuss imminent challenges and opportunities.

1. Introduction

Polymer informatics applies AI- and data-centric methods to search polymers’ vast chemical space and design materials with targeted properties. Its ecosystem spans data, representations, surrogate prediction, candidate validation, and synthesis planning, while facing major data and representation challenges.

  • Polymers occupy an effectively vast chemical space whose properties can be tuned through chemical and morphological structure.
  • Polymer informatics uses modern data- and information-centric approaches inspired by AI and ML to address polymer search.
  • Polymer data comes from experiments, high-throughput computations, literature extraction, and emerging autonomous computational agents, but organized data remains limited.
  • Machine-readable representations must encode polymer chemical and morphological information as well as synthesis and processing conditions.
  • Surrogate models use polymer fingerprints and target-property data to predict properties and design candidates, which then require computational or physical synthesis validation.
  • The paper reviews ecosystem components from data management and representation through prediction, design, applications, and critical next steps.

2. Data generation, acquisition and management

Polymer data are collected from literature, databases, experiments, and computations, with autonomous agents and hypothetical-polymer generation extending discovery beyond known materials. Major constraints include laborious extraction, complex multiscale structures, and difficult computation of diverse polymers.

  • Polymer data resources include handbooks, online repositories, journal articles, experiments, and high-throughput computations.
  • Manual literature excerption requires laborious extraction, validation, and normalization because polymer data lack standard publication policies.
  • Machine-learning NLP can scan literature and organize extracted polymer properties into tabular, machine-readable relations, but technical interpretation remains difficult.
  • High-throughput DFT and classical MD generate polymer data, but realistic amorphous and semicrystalline structures create substantial computational challenges.
  • Computing properties across diverse, understudied polymers remains challenging, particularly for properties modeled with classical molecular dynamics.
  • Autonomous computational agents iteratively select candidates, predict target properties, and convert promising polymers into hierarchical 3D structural models.
  • BRICS-based generation uses building blocks extracted from approximately 12,000 successfully synthesized polymers to explore realistic hypothetical polymers.

3. Polymer representation

Polymer informatics transforms polymer chemistry and morphology into machine-readable representations for property modeling. Existing approaches include oligomer SMILES, hierarchical fingerprints, BigSMILES, and graph-based methods, each with important coverage or scalability limitations.

  • Polymer representations transform collected structural, chemical, property, and synthesis data into formats suitable for AI and ML methods.
  • Group-contribution methods estimate polymer properties as weighted sums of contributions from constituent fragments.
  • SMILES-based oligomer fingerprints can predict polymer properties fairly well but exclude morphology effects.
  • Hierarchical polymer fingerprints encode atomic-, block-, and chain-level connectivity and morphology features for property prediction.
  • BigSMILES extends SMILES to homo- and copolymers and can incorporate branching, network, and terminal-group information.
  • Graph neural networks remain limited for polymers because large-scale morphology and repeat-unit connection points are difficult to represent.

4. Property prediction schemes

Polymer property prediction combines regression, probabilistic, multi-fidelity, and neural-network methods with polymer fingerprints to estimate diverse properties. Multi-fidelity models can improve accuracy with limited high-fidelity data, while dataset size, fidelity alignment, and computational cost remain important constraints.

  • Learning algorithms: Learning algorithms for polymer properties include linear and non-linear regression, multi-fidelity information fusion, and deep neural networks.The appropriate method depends on target-property complexity and the volume and nature of available data.
  • Linear/Non-linear regression: Gaussian process regression provides property predictions together with uncertainty estimates from its learned probabilistic distribution.Representative applications include glass transition temperature, chain band gap, frequency-dependent dielectric constant, and gas permeability.
  • Multi-fidelity information fusion approaches: Multi-fidelity models combine lower-cost, lower-accuracy data with precise high-fidelity data to improve property predictions, especially when high-fidelity datasets are small.The approach consolidates information across fidelity levels rather than relying only on the most expensive measurements or computations.
  • Multi-fidelity information fusion approaches: MF models surpassed single-fidelity GPR accuracy with substantially fewer high-fidelity samples in examples involving crystallization tendency and band gap.The examples used 107 high-fidelity and 429 low-fidelity crystallization data points, and 382 computed band-gap values across fidelity levels.
  • Multi-fidelity information fusion approaches: MF accuracy depends on learning the shared latent space between fidelity levels, while multiple fidelity hierarchies can increase model parameters and computational cost.The review identifies faster MF algorithms as necessary when polymer datasets contain several simultaneous fidelity levels.
  • Deep neural networks: Deep neural networks have been applied to solvent selection and to predictions of glass transition temperature, gas permeability, and thermal conductivity.The solvent model classified good-solvent and non-solvent polymer pairs using polymer and solvent descriptors, covering 4,595 polymers and 24 solvents.

5. Polymer design algorithms

Polymer design algorithms use surrogate models to screen enumerated candidates, guide experiments through active learning, and generate polymers directly through evolutionary or generative approaches. These strategies have identified candidates with application-relevant property combinations, while differing in search space, interpretability, and parameter requirements.

  • Enumeration: Surrogate property models rank large, pre-determined polymer pools against application-specific screening criteria.Predicted properties can be combined to identify candidates meeting multiple target requirements.
  • Enumeration: Machine-learning screening identified polymer dielectrics, high-temperature polymers, thermally conductive polymers, and gas-separation membranes with targeted properties.Reported examples include thermal conductivities of 0.18–0.41 W/mK and two membranes discovered from over 11,000 homopolymers.
  • Sequential (Active) learning: Active learning combines a surrogate model, an acquisition function, and iterative incorporation of new experiments into the knowledge dataset.Gaussian-process methods can provide both predicted property values and uncertainty estimates for acquisition decisions.
  • Sequential (Active) learning: 30, 46, 98 and 234 experiments were required to discover 10 polymers with Tg > 450 K using balanced, exploitation, exploration and random acquisition strategies, respectively.The averages were calculated across 50 runs, with standard deviations shown as error bars; balanced exploitation and exploration performed best.
  • Evolutionary and generative strategies: Inverse design directly generates polymers satisfying property objectives rather than screening only a predefined candidate set.Evolutionary algorithms use crossover, mutation and selection, while VAEs and GANs search or map through continuous latent spaces.
  • Evolutionary and generative strategies: Fast surrogate evaluation enables genetic algorithms to explore broad polymer spaces and optimize multiple property criteria simultaneously.The approach was applied to high-glass-transition-temperature, large-band-gap polymers, a difficult target combination found in only 4 of thousands of known polymers.
  • Evolutionary and generative strategies: Genetic algorithms are more interpretable and easier to tune than VAEs, whereas VAEs explore wider chemical spaces; comparative studies remain necessary to establish scenario-specific choices.Prior knowledge is also more readily incorporated into genetic algorithms through population and mutation biases.

6. Application examples

Polymer informatics is applied to diverse material challenges, including dielectric capacitors, gas separation, batteries, conducting polymers, bioplastics, and depolymerizable materials. Each application requires screening or design against distinct combinations of performance, stability, transport, and sustainability-related properties.

  • Polymer dielectrics: Polymer dielectric capacitors require combinations such as high dielectric constant and high breakdown strength for high-energy-density storage.Machine learning can screen candidates using glass transition temperature, dielectric constant, band gap and charge-injection barriers for extreme conditions.
  • Polymer dielectrics: Known and hypothetical-polymer enumeration, genetic algorithms and variational autoencoders have all been used to propose polymer dielectrics for high-temperature capacitors.Several computation- and data-driven strategies have produced designed and synthesized dielectric films with high dielectric constant and band gap.
  • Gas separation: Gas-separation membranes require high permeability and selectivity, but current membranes face low selectivity and physical aging.These constraints make it difficult to identify polymers that jointly improve the relevant transport properties.
  • Battery electrolytes: Solid polymer electrolytes for Li-ion batteries require a wide electrochemical stability window, high ionic conductivity and Li-ion transference, and low glass transition temperature.
  • Conducting polymers: Conducting-polymer discovery is complicated by electron-transfer mechanisms and interactions among dopants and polymers.Molecular doping can increase conductivity but also makes optimal polymer–dopant pair discovery more difficult.
  • Bioplastics: Bioplastic development requires structure–property relationships that identify application-specific chemistries compatible with nonconventional biosynthesis routes.The relevant feedstocks include highly oxygenated molecules derived from plants and bacteria.
  • Depolymerizable polymers: Depolymerizable polymers are sought for applications including drug delivery, recyclable plastics, self-healing and recyclable coatings, with low ceiling temperatures being desirable.Specific stimuli can trigger rapid depolymerization into monomers at moderate to relatively low temperatures.

7. Critical next steps

Critical next steps center on expanding polymer data and representations beyond simple chemical structures, while improving learning, synthesis planning, and autonomous experimentation for polymers' complex compositions and processing.

  • 7.1. Beyond homopolymers: Co-polymers, blends, additives, and nanocomposites remain largely unexplored because their chemical and physical structures are complicated and available data are sparse.Monomer or polymer ratios and structural arrangements significantly affect properties, but such data are difficult to collect systematically and dynamically.
  • 7.2. Sustainable data capture: Polymer informatics needs broad-based data infrastructure that captures untapped information from journal text, tables, and figures.Existing chemical-information extraction toolkits target inorganics and molecules; comparable tools for polymers still need development.
  • 7.3. Polymer representation and learning: Accurate universal property models must incorporate molecular weight, morphology, and processing conditions in addition to chemical fingerprints.These descriptors strongly affect crystallinity, mechanical properties, and solution behavior, whereas fingerprints already provide acceptable accuracy for several structure-dominated properties.
  • 7.3. Polymer representation and learning: Graph neural networks could learn task-specific polymer representations, but large polymer graphs and repeat-unit connection points remain difficult to model.Replacing polymers with oligomers is a potential solution, although its prediction capability still needs testing.
  • 7.4. Polymer synthesis planning: Polymer realization remains slow because synthesis requires identifying reactions, precursors, reagents, processing conditions, and practical constraints such as toxicity or cost.Synthesis pathways have historically depended heavily on domain knowledge and personal experience, motivating AI-assisted planning and robotic or autonomous retrosynthesis.
  • 7.5. Autonomous integration of experimental and computational workflows: Autonomous polymer design must address strong processing sensitivity of molecular weight and difficult real-time characterization of complex polymer structures.These challenges arise from sensitivity to processing time and conditions, semi-crystalline or amorphous structures, branching, and stereochemical relationships.
Loading 2011.00508v1…