Source-linked AI summary
Computer-aided molecular design: An introduction and review of tools, applications, and solution techniques
Nick D. Austin, Nikolaos V. Sahinidis, Daniel W. Trahan
TL;DR
CAMD addresses the difficulty of selecting suitable molecules within an enormous chemical design space and the limitations of trial-and-error design. This review presents QSPR-based descriptors, optimization formulations, constraints, solution techniques, and applications, concluding with a structured account of CAMD methods and design problems.
Problem
Chemical design must search spaces far larger than those manageable through trial-and-error, including more than 115 million cataloged structures and an even larger theoretical space.
Method
The article reviews three QSPR classes, CAMD formulations and constraints, and mathematical optimization, decomposition, heuristic, and generate-and-test solution approaches.
Results
The review consolidates QSPRs, formulations, feasibility constraints, solution techniques, and applications across single-molecule, mixture, and integrated product/process CAMD problems.
Takeaways & Limitations
CAMD provides a framework for exploring molecular structures through property models and optimization while addressing design constraints and chemical feasibility.
Takeaways & Limitations
Topological-index CAMD can face combinatorial difficulties and has so far been demonstrated only on small design problems.
Abstract
from arXiv · showhide
This article provides an introduction to and review of the field of computer-aided molecular design (CAMD). It is intended to be approachable for the absolute beginner as well as useful to the seasoned CAMD practitioner. We begin by discussing various quantitative structure-property relationships (QSPRs) which have been demonstrated to work well with CAMD problems. The methods discussed in this article are (1) group contribution methods, (2) topological indices, and (3) signature descriptors. Next, we present general optimization formulations for various forms of the CAMD problem. Common design constraints are discussed and structural feasibility constraints are provided for the three types of QSPRs addressed. We then detail useful techniques for approaching CAMD optimization problems, including decomposition methods, heuristic approaches, and mathematical programming strategies. Finally, we discuss many applications that have been addressed using CAMD.
1 Introduction
Chemical product design has traditionally explored only small regions of an enormous molecular design space, motivating computer-aided molecular design (CAMD). CAMD combines quantitative structure-property relationships with numerical optimization, and this review organizes its descriptors, formulations, solution techniques, and applications.
- Motivation: Chemical product design has often relied on laborious trial-and-error and small sets of related compounds, limiting exploration of possible molecular structures.Computational resources make broader design-space exploration more feasible.
- CAMD approach: CAMD uses semi-empirical quantitative structure-property relationships together with fast numerical optimization to search for suitable molecular structures.It combines molecular modeling, thermodynamics, and numerical optimization to design good or optimal molecules.
- Review scope: The review introduces group contribution methods, topological indices, and signature descriptors as three popular classes of QSPRs used in CAMD.These methods provide ways to relate molecular structures to properties at different levels of accuracy.
- Review scope: It presents CAMD formulations, design and structural-feasibility constraints, and solution techniques including mathematical optimization, decomposition, and heuristic approaches.The formulations cover multiple CAMD problem types and aim to ensure practical solutions and chemically feasible structures.
- Applications: The article concludes with a diverse, non-exhaustive review of CAMD applications for readers ranging from beginners to advanced practitioners.Its stated purpose is both to motivate CAMD from the ground up and to provide detailed tools and techniques for experienced users.
2 Popular types of QSPRs in CAMD
CAMD uses semi-empirical QSPRs to map molecular structure into quantitative properties and make large chemical design spaces tractable for optimization. The section reviews group contribution methods, topological indices, and signature descriptors, along with their strengths and limitations.
- Semi-empirical QSPRs connect abstract molecular structures to quantitative properties and enable efficient estimation for CAMD optimization.They represent molecules through substructures of atoms and bonds, allowing combinatorial optimization over molecular representations.
- Group-contribution methods: Group contribution methods predict properties from counts of molecular substructures, such as −CH3, −CH2−, and −OH in propanol.A coefficient vector is regressed from molecular data, while the group-count vector records each group’s occurrences.
- Group-contribution methods: Interaction terms and multilevel groups extend group contribution methods to capture simultaneous group effects and proximity effects within molecules.These extensions address cases where isolated group contributions do not adequately represent molecular-scale or neighboring-group effects.
- Group-contribution methods: Group contribution methods are intuitive, cover diverse structures through group combinations, and translate readily into mathematical optimization formulations.Their main limitations include weak isomer discrimination, inconsistent group sets across properties, and dependence on groups specified before regression.
- Topological indices and signature descriptors: Topological-index QSPRs describe molecules through graph-theoretic properties, whereas signature descriptors can be manipulated to represent both group contributions and topological indices.Topological-index design may cover only a particular chemical-space subset, be difficult to interpret chemically, and create combinatorial difficulties in CAMD.
3 CAMD as an optimization problem
CAMD formulates molecular design as a reverse problem: predicting structures from desired properties while handling large search spaces, structural feasibility, and consistency constraints. Mathematical optimization combines QSPR estimates with design, process, and feasibility constraints and an objective function.
- QSPR techniques estimate molecular properties from structural attributes, whereas CAMD reverses this direction by predicting structures from properties.
- CAMD optimization must search broadly while ensuring that candidate descriptor vectors correspond to chemically feasible and internally consistent molecular structures.
- The formulation uses QSPR functions to estimate properties from group counts, graph features, or atomic signatures.
- Constraint functions represent bounds on properties, structural features, process conditions, thermodynamics, cost, and system-specific interactions.
- Structural-feasibility functions test whether the descriptor vector is consistent with a molecular structure that can actually exist.
- The objective function quantifies molecular performance using properties and possibly descriptors, and may be minimized or maximized depending on the design problem.
3.1 Classes of the CAMD problem
CAMD formulates molecular design as selecting feasible structures for performance, property targets, or process-linked objectives. The section distinguishes single-molecule, mixture, and integrated process/product formulations.
- Single molecule design: Single-molecule CAMD covers finding all feasible structures, optimizing molecular performance, and matching properties to target values.Feasibility formulations can produce candidates later evaluated using higher-order models or experiments.
- Single molecule design: Exact performance functions can make the CAMD solution optimal under the supplied application-specific ranking criterion.An algebraic relationship between descriptors and performance also supports numerical optimization strategies.
- Single molecule design: Property-target formulations minimize distance from desired values and can support substitute molecules motivated by environmental, cost, or availability concerns.The common objective uses weighted squared distances between estimated properties and targets.
- Mixture design: Mixture design jointly represents component descriptors, mixture properties, thermodynamics, mole fractions, and non-ideal behavior.The formulation includes mixture-property relationships, mole-fraction normalization, and bounds on mixture properties.
- Integrated process and product design: Integrated process/product design adds process variables and process-related constraints to molecular design objectives.These problems are challenging because structure descriptors and process performance lack an easily discernible algebraic relationship.
3.2 Common design features in the form of constraints
CAMD constraints restrict descriptors, properties, process conditions, and mixture behavior so proposed structures remain within the intended design space and application requirements.
- Application-specific constraints: Constraint sets vary across CAMD applications because they encode process conditions, design necessities, thermodynamic requirements, and other problem-specific features.The review therefore summarizes recurring design features rather than detailing every application-specific constraint.
- Descriptor and property constraints: Descriptor-count constraints bound the number or total amount of structural features in a designed molecule.Lower bounds can ensure that more than one descriptor appears in the solution.
- Descriptor and property constraints: Fixed-descriptor constraints force selected groups or structural features to occur when designing analogues within a chemical family.The required descriptors are specified through a fixed descriptor set and occurrence bounds.
- Descriptor and property constraints: Property bounds prevent violations of process criteria, environmental regulations, and toxicity thresholds.These bounds are applied analogously to structural descriptor constraints.
- Mixture and thermodynamic constraints: Mixture CAMD commonly requires activity-coefficient models to represent species activities and non-ideal mixture behavior.UNIFAC, SAFT, and COSMO-based methods are identified as models used in CAMD-related thermodynamic calculations.
3.3 Forms of the structural feasibility constraints
Structural feasibility constraints connect QSPR variables to chemically realizable molecular graphs. Group-contribution, topological-index, and signature-descriptor methods require different consistency and connectivity conditions.
- Forms of the structural feasibility constraints: Structural feasibility constraints ensure selected descriptors are mutually consistent and can be assembled into a chemically feasible molecule.The required constraints differ among group-contribution, topological-index, and signature-descriptor QSPRs.
- 3.3.1 Connectivity constraints for GC methods: Group-contribution constraints use group counts and valences to enforce connectivity, with m distinguishing acyclic, monocyclic, and bicyclic compounds.Valence Φd denotes the number of external bonds required by group d.
- 3.3.1 Connectivity constraints for GC methods: Additional group-contribution constraints prevent valence-compatible but unjoinable group sets by requiring enough groups to satisfy every group’s valence.The formulation is extended for aromatic and aliphatic rings, with assumptions about benzylic attachment and bond types.
- 3.3.2 Connectivity constraints for TI methods: Topological-index formulations represent molecular connectivity with adjacency matrices whose entries encode bond types between vertices.Node valences and assignment constraints complement the adjacency representation.
- 3.3.2 Connectivity constraints for TI methods: Topological-index connectivity constraints require every non-dummy vertex to attach through a lower-index vertex, preventing disjoint subgraphs.This condition enforces a connected molecular graph.
- 3.3.3 Connectivity constraints for SD methods: Signature-descriptor constraints maintain consistency among overlapping descriptors by matching parent-child coloring relationships across signatures.The approach defines descriptor subsets by external bonds and uses graph-degree relationships based on the handshaking lemma.
- 3.3.3 Connectivity constraints for SD methods: Signature-descriptor feasibility also links edge counts, vertices, and cycle-space dimension, requiring balanced coloring sequences and connectable roots and children.The graph constraint need only be checked for one signature height when multiple heights are used.
4.1 Generate-and-test methods
Generate-and-test methods enumerate candidate molecular structures and evaluate them with QSPR models. They are intuitive and practical for reduced design spaces but inefficient when the space is large.
- Generate-and-test methods: QSPR models can evaluate millions of candidate structures within minutes, making forward generate-and-test screening computationally attractive.This approach is especially advantageous after the candidate pool has been reduced to a practical size.
- Generate-and-test methods: Large design spaces are not solved efficiently by generate-and-test, giving optimization methods a distinct advantage in those cases.Knowledge-based reduction procedures or smaller problems can make enumeration more practical.
- Generate-and-test methods: Generate-and-test algorithms require considering every possible structure while minimizing redundancy, with methods available for groups, topological indices, and signature descriptors.Their efficiency depends on the size of the candidate space.
4.2 Decomposition methods
Decomposition methods make difficult CAMD problems tractable by splitting them into smaller subproblems and progressively applying constraints. Their effectiveness depends on reducing the feasible molecular set before evaluating difficult constraints or objectives.
- 4.2 Decomposition methods: CAMD problems with large descriptor spaces or challenging nonlinear process models are often solved as successive optimization subproblems.Decomposition addresses problems that are too difficult to optimize directly.
- 4.2 Decomposition methods: Single-molecule decomposition typically applies structural feasibility constraints first, then evaluates remaining constraints and the objective on the resulting structures.Feasible descriptor vectors can be assembled into structures for one-by-one evaluation or further optimization.
- 4.2 Decomposition methods: Decomposition works best when few descriptors, tight constraints, or distance-to-target objectives reduce the feasible structure set to a manageable size.When many descriptor vectors remain, decomposition should be paired with another approach.
- 4.2 Decomposition methods: Mixture design can be decomposed into several single-component molecular design problems, whose solutions are then considered as potential mixture components.This strategy relies on the relative efficiency of single-molecule design.
- 4.2 Decomposition methods: Integrated product/process design can decompose by optimizing process variables with unrestricted product properties, producing ideal property targets for a subsequent molecular design problem.This relaxation gives a lower bound for minimization problems.
- 4.2 Decomposition methods: Alternative decompositions optimize molecular structures first, iteratively alternate process and molecular subproblems, or use outer approximation to generate candidate molecules.These approaches include Pareto-based structure screening and fixed-molecule process subproblems.
4.3 Mathematical optimization methods
Mathematical optimization methods provide formulations and algorithms for CAMD problems, especially when descriptor spaces are large or molecular and process models are nonlinear or nonconvex. The reviewed strategies include generate-and-test, mixed-integer programming, outer approximation, and branch-and-bound.
- 4.3 Mathematical optimization methods: Optimization is most suitable for CAMD problems with many possible descriptors or difficult nonlinearities and nonconvexities; smaller descriptor spaces can often be enumerated.Generate-and-test methods can efficiently handle problems with few possible descriptors.
- 4.3 Mathematical optimization methods: Single-molecule feasibility formulations can be solved as mixed-integer nonlinear programs using outer-approximation algorithms.These formulations have been applied to molecular design and structure connectivity.
- 4.3 Mathematical optimization methods: Branch-and-bound uses underestimators to obtain lower bounds, branches on variables, and discards search regions that cannot lead to optimal solutions.Upon convergence, it guarantees a globally optimal solution within a user-specified tolerance.
- 4.3 Mathematical optimization methods: Property-target design can be formulated as a mixed-integer quadratic program or transformed into a mixed-integer linear program.The formulation uses positive and negative deviation variables between target and estimated property values.
- 4.3 Mathematical optimization methods: A decomposition workflow identifies optimal group sets, generates structures, and applies higher-order models to feasible structures.Later extensions incorporate connectivity through additional variables.
4.4 Heuristics
Heuristic methods search complex CAMD spaces by generating and evaluating trial molecular structures according to high-level selection rules. Their success depends partly on how molecular structures are encoded and includes genetic, tabu-search, ant-colony, and simulated-annealing approaches.
- 4.4 Heuristics: Mathematical optimization may be impractical for very high-dimensional or highly nonlinear CAMD problems involving diverse molecules or difficult process simulations.Such problems may require difficult subproblems for each feasible structure.
- 4.4 Heuristics: Heuristics generate trial points, evaluate their objectives, and select subsequent points using objective values and the history of evaluations.Termination occurs at convergence or a user-defined time or iteration limit.
- 4.4 Heuristics: Molecular encoding is a central heuristic-design choice, ranging from descriptor counts to binary strings or descriptor-based values.The encoding determines how molecular structures are represented during search.
- 4.4 Heuristics: Genetic algorithms evaluate generations of candidate solutions and recombine parent representations with selection favoring better-ranked solutions.CAMD implementations encoded structures using substituent groups, UNIFAC groups, or other group identities.
- 4.4 Heuristics: Tabu search repeatedly modifies candidate molecules while forbidding solutions in a tabu list based on factors such as recurrence, infeasibility, or poor objective values.It has been applied to transition-metal catalyst and ionic-liquid design.
- 4.4 Heuristics: Ant-colony optimization attracts candidate solutions toward better solutions, whereas simulated annealing can accept worse solutions early before becoming more stringent.Both methods use iterative changes to encoded molecular descriptors.
5 Applications of CAMD
CAMD has been applied across single-molecule, mixture, and integrated product/process design problems. Applications include solvents, refrigerants, polymers, pharmaceuticals, ionic liquids, separations, and other industrial products and processes.
- Single-molecule design: Single-molecule CAMD applications include feasibility design, explicit-objective design, solvents, pharmaceuticals, refrigerants, and polymers.Examples target solvent performance, drug modification, refrigerant properties, and polymer properties such as glass transition temperature and density.
- Industrial processes: Industrial solvent design has addressed liquid-liquid extraction, extractive distillation, and reaction-property optimization.Reported applications include water–acetic acid separation, toluene replacement, ethanol fermentation, and reaction-rate optimization.
- Ionic liquids: Ionic-liquid CAMD has targeted electrical conductivity, heat transfer, liquid-liquid separations, solubility, and related properties.Both decomposition and tabu-search methods have been used for ionic-liquid design.
- Mixture design: Mixture design is more difficult than single-molecule design and consequently has fewer application examples, although many techniques are described as generalizable.Applications include polymer blends, cleaning-agent solvent mixtures, crystallization solvents, and paint or insect-repellent blends.
- Integrated product/process design: Integrated product/process design addresses sensitivity between product descriptors and process variables by considering both design problems simultaneously.Examples include carbon capture, carbon dioxide–methane separation, and fluids for organic Rankine cycles.
- Application summary: Table 3 summarizes CAMD applications together with the methodologies used in each case.The table organizes examples by application and method categories.
6 Conclusions
CAMD combines accurate QSPRs with efficient optimization to search vastly broader molecular design spaces than traditional trial-and-error approaches. The article reviews core QSPRs, optimization formulations, solution techniques, and applications supporting this capability.
- Computational resources, optimization algorithms, and QSPRs transformed chemical-product design into a rapid search through millions of molecular structures.This expanded the diversity of structures that designers could investigate beyond the small design spaces typical of trial-and-error practice.
- The review explains three QSPR families—group contribution methods, topological indices, and signature descriptors—and their associated feasibility constraints.
- It presents mathematical formulations for single-molecule, mixture, and integrated product/process CAMD problems.
- Decomposition methods, heuristic approaches, and mathematical programming strategies are discussed as techniques for solving difficult CAMD optimization problems.
- CAMD has supported improved industrial processes, new consumer products, and high-impact chemical-process optimization, with further applications emerging.