Source-linked AI summary

The Impact of Large Language Models on Scientific Discovery: a Preliminary Study using GPT-4

Microsoft Research AI4Science, Microsoft Azure Quantum

arXiv:2311.07361v2cs.CLcs.AI

TL;DR

This report examines GPT-4’s capabilities and potential limitations in selected natural-science domains. Using a broad evaluation centered on scientific applications, it finds considerable potential in drug discovery, including broad knowledge and novel molecule generation.

  • Problem

    The report asks how GPT-4 can support natural-science research across selected domains and what limitations constrain its scientific use.

  • Method

    The study evaluates GPT-4 across selected scientific tasks through broad assessments, including drug-discovery applications and case-based analysis.

  • Results

    GPT-4 shows considerable potential for drug discovery, with broad knowledge across key concepts and the ability to generate novel molecules from text instructions.

  • Takeaways & Limitations

    GPT-4 may provide useful insights, suggestions, and candidate molecules across a range of drug-discovery tasks.

  • Takeaways & Limitations

    The assessment relies substantially on case studies that are subjective, informal, and less rigorous than formal scientific evaluation.

Abstract

from arXiv · show

In recent years, groundbreaking advancements in natural language processing have culminated in the emergence of powerful large language models (LLMs), which have showcased remarkable capabilities across a vast array of domains, including the understanding, generation, and translation of natural language, and even tasks that extend beyond language processing. In this report, we delve into the performance of LLMs within the context of scientific discovery, focusing on GPT-4, the state-of-the-art language model. Our investigation spans a diverse range of scientific areas encompassing drug discovery, biology, computational chemistry (density functional theory (DFT) and molecular dynamics (MD)), materials design, and partial differential equations (PDE). Evaluating GPT-4 on scientific tasks is crucial for uncovering its potential across various research domains, validating its domain-specific expertise, accelerating scientific progress, optimizing resource allocation, guiding future model development, and fostering interdisciplinary research. Our exploration methodology primarily consists of expert-driven case assessments, which offer qualitative insights into the model's comprehension of intricate scientific concepts and relationships, and occasionally benchmark testing, which quantitatively evaluates the model's capacity to solve well-defined domain-specific problems. Our preliminary exploration indicates that GPT-4 exhibits promising potential for a variety of scientific applications, demonstrating its aptitude for handling complex problem-solving and knowledge integration tasks. Broadly speaking, we evaluate GPT-4's knowledge base, scientific understanding, scientific numerical calculation abilities, and various scientific prediction capabilities.

1 Introduction

This report examines GPT-4’s capabilities and potential limitations across selected natural-science fields, including drug discovery, biology, computational chemistry, materials design, and PDEs. It finds broad promise across these domains while noting accuracy and evaluation limitations.

  • Scope: The study focuses on GPT-4’s performance and applicability in drug discovery, biology, computational chemistry, materials design, and partial differential equations.These areas were selected because covering all natural-science sub-disciplines was infeasible.
  • Overall findings: GPT-4 demonstrates potential across scientific domains, with capabilities spanning diverse tasks and strong understanding of key concepts.The report evaluates knowledge, scientific understanding, numerical calculation, and prediction capabilities.
  • Domain findings: In drug discovery and biology, GPT-4 provides useful insights, predictions, and assistance across broad task ranges.Reported applications include drug-target binding, molecular properties, retrosynthesis, biological-language processing, bioinformatics, and biology design.
  • Domain findings: In computational chemistry and materials design, GPT-4 can retrieve information, suggest methods, generate code, and propose designs, but struggles with accurate coordinates and quantitative predictions.The report identifies challenges with complex molecular or material structures and precise numerical outputs.
  • Domain findings: For PDEs, GPT-4 can explain concepts, recommend analytical and numerical methods, generate code, and provide proof approaches, while theorem proving remains limited.Its independent discovery and validation capabilities also require improvement.
  • Limitations: A major limitation is that much of the assessment relies on subjective, informal case studies rather than rigorous formal evaluation.The authors call for more formal and comprehensive testing methods.

2 Drug Discovery

GPT-4 shows broad knowledge and versatility across drug-discovery tasks, but its reliability varies substantially by task. It performs well on qualitative knowledge and some similarity-based interaction prediction, while molecular representations, affinity prediction, and quantitative calculations remain challenging.

  • 2.2 Understanding key concepts in drug discovery: GPT-4 demonstrates broad knowledge of drugs, target proteins, physicochemical principles, and challenges across the drug-discovery process.The report describes useful responses about Afatinib, 3CLpro, Lipinski’s Rule of Five, and broader drug-discovery challenges.
  • 2 Drug Discovery: GPT-4 supports molecule manipulation, drug-target interaction prediction, molecular property prediction, retrosynthesis, novel molecule generation, and coding for drug discovery.These capabilities are presented as potentially useful across candidate generation, optimization, prediction, synthesis planning, and computational workflows.
  • 2.1 Summary: GPT-4’s quantitative-task limitations include inaccurate molecular calculations and predictions, motivating further improvement such as fine-tuning.The report contrasts GPT-4’s qualitative strengths with weaker accuracy on quantitative tasks.
  • 2.2.1 Entity translation: IUPAC translation is more reliable than SMILES generation: GPT-4 correctly translated SMILES to IUPAC but failed in the reverse direction and generated incorrect formulas in both directions.For Afatinib, GPT-4 also produced the correct chemical formula and IUPAC name but an incorrect SMILES sequence.
  • 2.3 Drug-target prediction: GPT-4 performs inconsistently on drug-target prediction: it appears random on BindingDB Ki, remains below deep-learning models on DAVIS, and improves strongly with embedding-based kNN examples for DTI.In the kNN DTI evaluation, increasing neighbors from k = 1 to k = 20 significantly improves accuracy, precision, recall, and F1, slightly surpassing BridgeDTI.
  • 2.3.1 Drug-target affinity prediction: Similarity-based few-shot examples improve drug-target affinity prediction, with Pearson Correlation approaching 0.5, although performance remains behind existing models.More similar examples further improve performance, with an upper bound observed when providing 30 nearest neighbors.

3 Biology

GPT-4 shows substantial potential for biological information processing, reasoning, and design assistance, but its accuracy is limited by sequence handling, quantitative tasks, and context-sensitive errors.

  • Biological information processing: GPT-4 processes specialized biological files, performs bioinformatic analyses, and predicts signaling peptides from provided sequences.Examples include MEME, FASTQ, and VCF formats, alongside signaling-peptide prediction.
  • Biological understanding and reasoning: GPT-4 demonstrates broad biological understanding and can reason about plausible mechanisms using built-in biological knowledge.The reported areas include consensus sequences, protein-protein interactions, signaling pathways, evolutionary concepts, and observations requiring mechanistic reasoning.
  • Sequence-related tasks: GPT-4 handles some sequence-related tasks successfully, including all-correct MYC binding-site predictions and signaling-peptide identification, but gives inconsistent ZNF143 predictions and incorrect references.Prompt changes produced different ZNF143 consensus sequences, whereas candidate-order shuffling did not affect MYC predictions.
  • Limitations: Direct biological sequence processing can cause catastrophic errors, while quantitative calculations, Arabic numerals, under-studied entities, and prompt phrasing also reduce reliability.The authors recommend manual verification, alternative computational tools, text-converted numerals, and refined prompting.
  • Biological assistance and design: GPT-4 assists biology design by estimating DNA properties, designing sequences, and translating experimental protocols into robot-control code.The reported design tasks include molecular computation and automatic pipetting workflows.
  • Limitations: GPT-4 may produce technically correct DNA designs that perform poorly because it does not account for flexible DNA strands and application context.The example uses Hamming-distance orthogonality without considering strand flexibility or related application requirements.

4 Computational Chemistry

GPT-4 shows broad potential as an assistant in computational chemistry, supporting literature review, method selection, simulation setup, coding, and research guidance. Its performance remains limited for complex atomic-coordinate generation, precise calculations, and some quantum-chemistry derivations.

  • GPT-4 can assist computational-chemistry researchers through literature review, method and software selection, simulation setup, code development, and experimental, computational, and theoretical guidance.These capabilities span electronic-structure methods, molecular dynamics, and practical research workflows.
  • Simulation and implementation assistant: GPT-4 can generate simple molecular structures and input files, but it is not adept at raw atomic coordinates for complex molecules or materials.For silicon, iterative correction produced a valid Quantum Espresso input file, while predictions of the generated structure’s results failed; simple molecules such as CH4 are more tractable.
  • Understanding of quantum chemistry and physics: GPT-4 understands many quantum-chemistry concepts but can make logically incorrect claims about size extensivity, size consistency, and antisymmetry derivations.It correctly identifies several density-functional-theory concepts, yet misclassifies CIS and MP2 and makes an algebraic error involving wavefunction antisymmetry.
  • Understanding of quantum chemistry and physics: GPT-4 can verbally explain physics processes, but zero-shot graphic expression and quantitative calculation remain unreliable.It correctly defines Feynman diagrams and describes target processes verbally, yet cannot directly generate a correct example diagram and requires improvement for numerical tasks.
  • Simulation and implementation assistant: GPT-4 can provide reasonable computational workflows and code, including identifying PySCF’s lack of an internal MRCI implementation and generating a Hartree-Fock workflow.Its recommendation for reducing Hartree-Fock computational cost was nevertheless invalid.

5 Materials Design

GPT-4 shows strong performance in materials-design knowledge retrieval, composition generation, synthesis planning, and coding assistance, but weaker quantitative and structure-generation performance. Its results vary by material class and task, with spatial reasoning and chemically valid representations remaining important limitations.

  • Knowledge memorization and design-principle summarization: GPT-4 excels at memorizing materials information and suggesting design principles, providing factual inorganic-crystal answers with only rare categorization mistakes.For solid-state electrolytes, 7 of 8 summarized design rules were judged correct, while one was factual but not a design principle.
  • Synthesis planning and coding assistance: GPT-4 provides satisfactory inorganic-material synthesis planning and generally helpful coding assistance for molecular-dynamics and DFT workflows.Generated code may still require iterative feedback and manual adjustment to refine inputs and processing pipelines.
  • Property prediction: GPT-4 predicts material properties poorly in quantitative settings, with metallic-versus-semiconducting classification only slightly better than a random guess.The report suggests additional training data, molecular graphs, or dedicated AI models as possible improvement routes.
  • Polymers and metal-organic frameworks: Performance differs across materials: polymer structure representations are difficult, while MOF reasoning often relies on connection-point or atom-count heuristics rather than full geometric analysis.GPT-4 failed to identify the highest-PLD MOF in all five experiments, although it selected the highest-PLD linker in two cases.
  • Composition creation: GPT-4 generates feasible inorganic compositions and performs well for alloys, perovskites, half-Heusler compounds, and spinels, but struggles with ternary and quaternary charge or element-count constraints.Alloy composition generation achieved a 100% success rate for producing the correct number of elements.
  • Structure generation: GPT-4’s atomic-structure capabilities are qualitatively useful but quantitatively limited: it correctly reported coordination environments for 34 of 84 examples, while generated crystal structures were usually unreasonable.Careful prompting and additional information such as space group and lattice parameters were needed to obtain more sensible structures.

Appendix C.5, GPT-4 does not perform well qualitatively as well.

GPT-4 shows some capability in materials-property prediction and synthesis assistance, but its quantitative accuracy and stable-structure generation remain limited. Performance varies by task, with qualitative predictions often stronger than precise numerical calculations.

  • Structure prediction: GPT-4 is unlikely to generate stable material structures from composition alone, while polymer atomic structures are especially difficult to predict.The report suggests using coding capabilities for more complex polymer-structure tasks.
  • Inorganic materials: MatBench evaluations cover metallic classification and electronic-band-gap regression using few-shot prompts, but prompt choice may affect results.The analysis uses the reported prompts and expects prompt variation to change quantitative results more than qualitative conclusions.
  • Inorganic materials: GPT-4 performs better than random guesses on MatBench tasks but remains far below state-of-the-art performance for inorganic-material property prediction.The results support some calculation and prediction capability, while dedicated models or further development are still needed for accuracy.
  • Polymer properties: For polymers, GPT-4 predicts qualitative thermal-conductivity properties reasonably but falls short on quantitative answers.A polymer glass-transition example identifies atactic polystyrene as having the highest Tg among the listed polymers.
  • Polymer properties: Increasing the number of few-shot examples generally improves polymer-property prediction, although it requires many demonstrations.Dielectric-constant prediction remains relatively inaccurate, with MAE and MSE values described as large.
  • Synthesis assistance: GPT-4 can assist synthesis planning and coding, but generated routes may omit reaction details, use imprecise conditions, or require feedback and guidance.The report describes correct prototypes and qualitatively correct steps in one example, while broader synthesis-route retrieval is mixed.

6 Partial Differential Equations

GPT-4 can support PDE education, concept analysis, solution-method selection, code generation, and research ideation across diverse PDE tasks. However, its analytical derivations and generated solutions can contain errors, so expert verification remains necessary.

  • 6.4 AI for PDEs: GPT-4 recommends analytical and numerical methods for PDEs, generates MATLAB and Python code, and proposes research directions such as extensions and generalizations.These capabilities span exact or approximate solution approaches, numerical implementation, and suggested new problems or improvements.
  • 6.2 Knowing basic concepts about PDEs: GPT-4 explains PDE definitions, classifications, applications, and specialized relationships, including distinctions among stochastic-PDE solution concepts and noise types.It describes mild and weak solutions as inclusively related and equivalent under specific conditions, while also discussing trace-class noise and space-time white noise.
  • 6.3 Solving PDEs: GPT-4 can provide accurate PDE solutions in some cases, including a self-similar solution whose result was accurate and not simply copied from the source.The derivation differed slightly from the book’s example while producing the correct result.
  • 6.3.1 Analytical solutions: GPT-4 incorrectly applies separation of variables to a non-homogeneous PDE, omits coupled terms, and may produce invalid analytical derivations.In one example, it treats coupled expressions as functions of separate variables and abandons T(t) while solving a resulting ODE.

7 Looking Forward

The study identifies strengths and limitations of GPT-4 for scientific discovery and discusses directions for improving or extending LLM-based systems. It emphasizes tool integration and unified scientific foundation models as promising paths forward.

  • GPT-4 shows proficiency across scientific tasks, including literature synthesis, property prediction, and code generation, but also produces inconsistent responses and occasional hallucinations.
  • Specialized datasets, architectures, and multi-task learning are proposed to improve SMILES, FASTA, and quantitative scientific task performance.
  • The authors argue that using LLMs alone is insufficient for scientific discovery and propose integrating them with scientific computation tools or building scientific foundation models.
  • Scientific tools and plugins are described as having potential to improve accuracy and reliability while helping researchers address complex problems.
  • A unified scientific foundation model should support multimodal, multiscale inputs spanning text, sequences, graphs, three-dimensional structures, and biomolecules.
  • Incorporating physical laws and first principles into model architecture and training is proposed because scientific data reflect observations governed by physical laws.

B Appendix of Computational Chemistry

The appendix evaluates GPT-4 on molecular-property prediction, molecular dynamics, Feynman-diagram drawing, and molecular-structure generation. Results indicate some useful physical reasoning and structural knowledge, but quantitative accuracy remains limited.

  • Molecular-property prediction: GPT-4 maps SMILES descriptions to molecular properties using the OGB and QM9 datasets, including the HOMO-LUMO gap and 12 QM9 properties.
  • Molecular-property prediction: For 11 QM9 properties, mean absolute errors decrease as more examples are provided, while GPT-4 gives a detailed but inaccurate calculation procedure for one property.
  • Feynman diagrams: GPT-4 is prompted to draw increasingly complicated Feynman diagrams, including cases with more informative system messages and reference diagrams.
  • Molecular structures: GPT-4 describes methane as one carbon atom surrounded symmetrically by four hydrogen atoms in a tetrahedral arrangement and provides xyz coordinates.
  • Molecular dynamics: The appendix compares absolute errors for MD17 energies and forces under random rotations using different numbers of provided examples.
  • Materials knowledge: For negative-Poisson-ratio materials, GPT-4 identifies the auxetic class and explains that stretching in one direction expands the perpendicular direction.

C.2 Knowledge memorization and design principle summarization for polymers

GPT-4 demonstrates substantial knowledge of polymer properties and common polymer names, but its polymer-structure representations and novelty claims can be unreliable. The examples include incorrect structures and limited novelty among proposed inorganic compounds.

  • Structure understanding: The bisphenol A response gives the correct formula C15H16O2 but presents an incorrect or confused structure involving a sorbitan ring.
  • Structure understanding: The Tween 80 response correctly identifies its common name and use, but the associated structure is evaluated as nonsense and confusingly combines Tween80 with sorbitan.
  • Polymer knowledge: GPT-4 has a clear understanding of polymer properties and recognizes common polymer names, but struggles to draw complex polymer structures in ASCII.
  • Candidate generation: Among 20 proposed inorganic solid-electrolyte materials, 3 are identified as novel, while many others are known materials or simple substitutions.

C.4 Representing polymer structures with BigSMILES

The BigSMILES cases show that GPT-4 can discuss polymer concepts and revise representations with guidance, but has minimal direct command of BigSMILES and makes structural errors. Correct representations emerge through iterative correction rather than reliable first-pass generation.

  • Nafion: The simplified Nafion representation treats the material as a random copolymer, with a perfluorocarbon backbone and sulfonic-acid side chains.
  • Atactic polypropylene: GPT-4 initially misinterprets the wildcard notation and omits a carbon in atactic polypropylene, before producing the closer representation {[$]C(C)C[$]}.
  • Random copolymers: For an ethylene-propylene random copolymer, corrections require linear ethylene notation and two connection points per repeat unit, yielding [$]CC[$],[$]C(C)C[$].
  • BigSMILES syntax: GPT-4 explains that separate connection descriptors mark independently connecting repeat units and commas indicate their random distribution.
  • BigSMILES knowledge: GPT-4 has minimal knowledge of BigSMILES notation and requires significant coaching to produce structures that still retain minor errors.
  • Candidate generation: GPT-4 can extract polymer concepts but cannot reliably propose specific novel polymer systems or represent complex polymers in SMILES formats.

C.5 Evaluating the capability of generating atomic coordinates and predicting structures using a novel crystal identified by crystal structure prediction.

GPT-4 gives qualitatively reasonable assessments of LiGaOS atomic structure, stability, and competing phases, but its quantitative predictions are poor. The evaluation reports especially broad uncertainty for ionic conductivity and a large energy-above-hull error.

  • GPT-4’s qualitative assessments of atomic structure, stability, and competing phases are mostly reasonable, but its quantitative evaluations are poor.The structural prediction lacks detailed atomic-resolution information, while ionic-conductivity estimates span a large range.
  • The case uses a novel crystal identified through crystal-structure prediction to assess GPT-4’s structure-generation and property-assessment capabilities.
  • GPT-4 was prompted to predict LiGaOS’s atomic structure, ionic-conductivity range, stability, synthesizability, energy above the hull, and competing phases.
  • 50 meV/atom versus 233 meV/atom is the reported GPT-4 estimate and ground truth for energy above the hull.The qualitative stability description and competing-phase analysis were nevertheless judged reasonable.

C.6 Property prediction for polymers

GPT-4 is evaluated on polymer-property comparison and prediction using qualitative judgments and experimental reference values. It generally provides thorough qualitative descriptions and can represent some quantitative properties, but quantitative thermal-conductivity prediction remains limited.

  • The evaluation asks which properties can compare polymer materials and which listed polymer has the highest Tg.
  • GPT-4’s statements about polymer properties were judged thorough and accurate in one property-prediction case.
  • GPT-4 represented quantitative and qualitative polymer properties well when compared with experimental glass-transition temperatures.Reference Tg values were approximately 183 K for 1,4-polybutadiene, 368 K for atactic polystyrene, and 291 K for PG-PPO-PG copolymers.
  • GPT-4 accurately predicted qualitative thermal-conductivity behavior for a novel two-dimensional crystalline C60 polymer but not its quantitative value.The comparison concerned the material’s thermal conductivity relative to molecular C60.

C.7 Evaluation of GPT-4 ’s capability on synthesis planning for novel inorganic materials

GPT-4 generates multiple suggested synthesis routes and conditions for novel inorganic materials. The examples include solid-state and hydrothermal approaches, but the proposed conditions are explicitly general and require experimental optimization.

  • GPT-4 proposes at least two synthesis routes for each requested inorganic compound, including solid-state and hydrothermal methods.
  • The solid-state route uses stoichiometric precursors, mixing and grinding, furnace heating, slow cooling, and product characterization.For Li0.388Ta0.238La0.475Cl3, the suggested atmosphere is inert gas or flowing dry HCl, at 600-800°C for 10-24 hours.
  • The hydrothermal route dissolves precursors in water, heats them in a Teflon-lined autoclave, then filters, washes, dries, and characterizes the precipitate.The suggested hydrothermal conditions are 180-240°C for 24-72 hours.
  • Example Ag2Mo2O7 routes use either Ag2O and MoO3 powders or AgNO3 and ammonium molybdate in aqueous processing.The powder route heats in air at 500-700°C for 10-24 hours; the precipitation route calcines at 400-600°C for 2-6 hours.
  • The suggested synthesis conditions may require optimization of precursors, heating rates, reaction times, and related parameters to achieve high phase purity.

C.8 Polymer synthesis

GPT-4 develops polymer-synthesis experiments using catalyst, temperature, and monomer-flow-rate combinations, then adapts the design to catalyst-dependent temperature ranges. It also proposes an information-seeking initial subset, while acknowledging that reactor-specific choices require adjustment.

  • Experimental design: The initial full-factorial design varies monomer flow rate, catalyst, and temperature across three levels, requiring 27 trials.Each trial measures the resulting polymer’s Young’s modulus, targeting isotactic polypropylene between 1350 and 1450 N/cm2.
  • Scope boundary: Catalyst choices and temperature ranges may need adjustment for the specific reactor setup and other factors, with literature and data used for fine-tuning.
  • Experimental design: The catalyst-dependent design uses 3 whole plots and 9 flow-rate/temperature subplots per catalyst, again totaling 3 × 9 = 27 trials.This split-plot structure accounts for temperatures depending on catalyst identity.
  • Information gain: The initial nine trials select every catalyst-flow-rate combination at each catalyst’s middle temperature to obtain preliminary information efficiently.Subsequent trials explore temperature effects for the most promising catalyst and flow-rate combinations.
  • Evaluation: GPT-4 is described as highly capable at experimental planning overall, despite losing track of the trial count in a more complicated final design.

C.9 Plotting stress vs. strain for several materials

GPT-4 generates plotting code for stress–strain curves and several materials-science relationships, but the examples expose omissions and code errors that limit reliability.

  • Stress–strain plotting: The stress–strain example appears reasonable but omits information about plastic deformation.
  • Stress–strain plotting: GPT-4 generates matplotlib code for stress–strain curves covering elastic and plastic deformation regions.The prompt requests several materials and explicitly includes permanent-deformation regions.
  • Semiconductor relationships: GPT-4 proposes plots relating band gap to lattice parameter, PBE band gap to experimental band gap, and band gap to alloy content.The alloy-content example uses AlxGa1-xAs, InxGa1-xAs, and InxAl1-xAs and illustrates band-bowing effects.
  • Plotting limitations: The band-gap–lattice-parameter plotting code cannot run, and the pressure–temperature plotting code contains an error.The pressure–temperature example uses water as the material system.
  • Semiconductor relationships: The alloy-content plot is designed to show band-gap bowing using Vegard’s Law and a band-gap bowing model.The implementation labels alloy content as x and band-gap energy in eV.

C.10 Prompts and evaluation pipelines of synthesizing route prediction of known inorganic materials

The synthesis-route evaluation prompts GPT-4 to predict inorganic-material precursors, reactions, and procedures, then scores these outputs against script-generated reference data. Results show correct precursor identification in one case but also expose errors caused by incomplete references and inaccurate procedures.

  • Prompting and data: The pipeline asks GPT-4 to predict a synthesis route from a compound’s common name and balanced chemical formula.The prompts frame GPT-4 as a materials-scientist assistant for synthesis tasks.
  • Prompting and data: Reference examples are script-generated plain-text records containing precursors, balanced reactions, and step-by-step synthesis procedures.The syntax is described as lackluster, but it represents information from the text-mining synthesis dataset.
  • Evaluation pipeline: The evaluation checks precursor-formula correctness and uses GPT-4 to assign normalized route-accuracy scores with or without explanations.The scoring prompts compare GPT-4-generated routes with script-generated reference routes.
  • Results: For one material, GPT-4 correctly predicts the precursors and balanced reaction but adds synthesis steps absent from the reference dataset.Those precursor and reaction predictions are characterized as trivially deducible from the product.
  • Results: For the same case, GPT-4 proposes solid-state sintering instead of the paper’s melt-quenching procedure, and its score misses the error because the reference entry is incomplete.The paper’s procedure melts the powders at 1250 °C for 15 minutes, whereas GPT-4 proposes 1300 °C for 12 hours.
  • Results: For Bi2MoO6, GPT-4 identifies only one precursor and proposes an incorrect balanced reaction, although its synthesis includes a paper-reported additional calcination step.Its 773 °C sintering temperature is close to the paper’s 700 °C value and differs from the reference dataset entry.

C.11 Evaluating candidate proposal for Metal-Organic frameworks (MOFs)

The MOF evaluation tests whether GPT-4 can match building blocks to topologies and design structures maximizing pore limiting diameter. The cases show that GPT-4 often relies on connection counts or atom counts without adequately reasoning about geometry.

  • Candidate compatibility: The first task asks GPT-4 to decide whether MOF building blocks and a topology can be assembled into a reasonable structure.The task requires spatial understanding of three-dimensional building blocks and topology compatibility and is based on PORMAKE.
  • Candidate compatibility: The preliminary study examines tbo for HKUST-1 and pcu for MOF-5, requiring two nodes for tbo and one node for pcu.The tbo net is 3,4-coordinated, whereas pcu is 6-coordinated.
  • Candidate compatibility: GPT-4 sometimes reaches the correct reject decision for MOFs while giving incorrect reasoning based on connection-point counts.It incorrectly claims pcu admits only 4-connected nodes and mischaracterizes the tbo topology’s node connectivities.
  • Candidate compatibility: Correct connection-point counts are insufficient to establish compatibility between a building block and a topology.The evaluation cases explicitly note that GPT-4 does not reliably account for the required structural compatibility.
  • PLD design: The design task selects among pcu MOFs to maximize pore limiting diameter using compatible metal nodes and sampled linkers.The candidate metal nodes have RMSD below 0.03 Å from the topology node local structure.
  • PLD design: GPT-4 failed to select the highest-PLD MOF in all five experiments, despite consistently choosing the metal node with the most atoms.Its selected ranks were 3rd, 6th, 15th, 3rd, and 11th/12th when sorted from high to low PLD.
  • PLD design: GPT-4’s PLD reasoning ranks larger metal nodes and longer, more flexible linkers as producing higher PLDs.The supplied example identifies the largest metal node and linker as the combination expected to yield the highest PLD in pcu.
  • PLD design: In design examples, GPT-4 considers building-block size through atom counts rather than geometry and reaches the wrong conclusion.
Loading 2311.07361v2…