Source-linked AI summary

Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study

Amelia Liu

arXiv:2608.14631v1cs.AI

TL;DR

The accuracy and reliability of LLMs for cosmetic chemistry and skin-health guidance remain poorly characterized despite growing consumer use. This study benchmarked 14 models across cosmetic chemistry and skincare tasks, finding strong retrieval performance but substantial difficulty with structural reasoning and algorithmic precision. The findings indicate that LLMs are not yet reliable sources for cosmetic chemistry or precision skincare guidance.

  • Problem

    The scientific accuracy and reliability of LLM-generated cosmetic chemistry and skin-health guidance remain poorly characterized despite growing consumer use.

  • Method

    The study benchmarked 14 LLMs from seven developers across quantitative cosmetic chemistry tasks and qualitative skincare scenarios.

  • Results

    LLMs performed well on familiar retrieval-based tasks but struggled significantly with multi-step structural reasoning, algorithmic precision, and synthetic-preservative identification.

  • Takeaways & Limitations

    LLMs are not yet reliable sources for cosmetic chemistry or precision skincare guidance, so consumers should treat their information as a starting point rather than a final answer.

  • Takeaways & Limitations

    Authoritative tone and plausible-sounding errors may lead consumers to place unwarranted trust in technically incorrect LLM responses.

Abstract

from arXiv · show

As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic chemistry, including the chemical properties of specific cosmetic ingredients and common cosmetic scenarios that may be of interest to consumers. Web search was disabled throughout to assess each model's internalized knowledge rather than its internet retrieval capacity. Overall performance was poor, with the most pronounced deficits in quantitative reasoning and structural identification tasks. While models handled general skincare questions with reasonability, responses consistently lacked the technical depth required for informed consumer decision-making. Notably, conversation with AI can pose a risk: outputs that sound authoritative but contain technical errors are less likely to generate skepticism compared to responses that explicitly acknowledge uncertainty. These findings suggest that general-purpose LLMs, trained predominantly on unverified public data, are currently not reliable sources of cosmetic chemistry information. Progress on two fronts, fine-tuning verified chemical and dermatological datasets, and substantial improvements to algorithmic reasoning, will likely be needed before these tools can be considered as resources for public use.

1 Introduction

LLMs are increasingly used for accessible skincare guidance, but their accuracy and reliability in specialized cosmetic chemistry and skin-health advice remain poorly characterized. This study benchmarks 14 popular LLMs across five topic areas, finding stronger performance on retrieval-based tasks than on multi-step chemical reasoning.

  • Context: LLMs have become integrated into daily life, including specialized inquiries and skincare guidance.Their applications span academic assignments, business automation, theoretical physics, and skin care guidance.
  • Motivation: Consumers often turn to online resources and LLMs for skincare information because professional dermatological consultation can be costly and time-intensive.Online resources vary considerably in quality and scientific rigor, while LLMs offer immediate availability and an interactive format.
  • Problem: The scientific accuracy and reliability of LLM responses in cosmetic chemistry remain poorly characterized, and their use for medical or technical advice is frequently cautioned against.The accessibility of conversational AI may benefit general consumers, but its suitability for sound skin-health guidance remains debated.
  • Related work: ChatGPT showed conversational competence but failed in rigorous general-chemistry problem-solving, while broader benchmarks found inconsistent performance in technical reasoning.These findings describe a gap between general explanation and rigorous, multi-step reasoning.
  • Study scope: The study evaluates 14 popular LLMs from seven leading AI developers across five cosmetic-chemistry and skin-health topic areas.The benchmark spans both quantitative tasks and qualitative skincare scenarios.
  • Study contribution: Models performed better on retrieval-based tasks than on tasks requiring multi-step chemical reasoning, regardless of parameter count.The results indicate both promise and current limitations in applying LLMs to cosmetic chemistry and skin health.

2 Experiment Design and Method

The study benchmarked 14 LLMs from seven developers using standardized cosmetic chemistry and skincare prompts spanning quantitative reasoning and qualitative consumer advice. Responses were collected without web access, independently validated, and analyzed for accuracy and parameter-count relationships.

  • Model Selection: Fourteen models from seven developers were tested, pairing flagship Pro or Large variants with efficiency-oriented Flash or Mini variants.The cohort covered Anthropic, DeepSeek, Google, Meta, Mistral AI, OpenAI, and xAI.
  • Evaluation Tasks: Five thematic inquiries elicited concise, one-sentence responses covering four quantitative chemistry tasks and one qualitative skincare-advice task.Quantitative topics addressed bonding, formulation preservatives, molecular weights, and Joback boiling-point estimates; qualitative prompts addressed sequential care and first-step recommendations.
  • Validation: All scenarios were predefined, based on frequently used cosmetic ingredients and products, and designed to simulate common consumer interactions.Quantitative answers were cross-referenced against NCBI PubChem.
  • Prompt Control: Queries were submitted independently with context retention disabled, reducing conversational drift and carryover guidance across responses.The standardized setup used a unified API gateway and automated submissions across models.
  • Knowledge Control: Web access was disabled so outputs reflected models’ pre-trained, parameter-based knowledge rather than real-time information retrieval.Responses therefore remained tied to each model’s frozen knowledge cutoff date.
  • Statistical Analysis: Accuracy was calculated as correct responses divided by queried ingredients or products, and parameter-count prediction was tested with linear regression using p < 0.05 for significance.Outputs were extracted programmatically with Python and organized into model-by-query accuracy tables.

3 Main Results

LLMs performed unevenly across cosmetic chemistry tasks: they were relatively reliable for molecular-weight calculations but struggled with bond counting, preservative identification, and Joback boiling-point estimation. For general skincare questions, responses showed recurring practical themes and frequent dermatologist disclaimers despite limited consensus in some scenarios.

  • Chemical structure and property tasks: Five models—the GPT-5 series, Grok-4 series, and Gemini 2.5 Pro—achieved perfect accuracy on all sigma- and pi-bond-counting stimuli.Scoring was all-or-nothing: both bond counts had to be precisely correct.
  • Chemical structure and property tasks: Mistral Large scored 0% accuracy and DeepSeek-R1 40% on sigma- and pi-bond counting.The disparity was discussed in relation to possible corpus and nomenclature biases across training datasets.
  • Chemical structure and property tasks: Preservative-identification accuracy was generally low, with no significant relationship between parameter count and performance (p = 0.34).A response counted as accurate only when all recognized synthetic preservatives in a product were correctly identified.
  • Chemical structure and property tasks: Six out of 14 models achieved 100% accuracy for molecular-weight calculations, while the remaining models correctly identified 4 out of 5 weights.This task outperformed synthetic-preservative identification, consistent with greater reliability for structured calculations based on widely available chemical data.
  • Chemical structure and property tasks: 7 out of 14 models failed to return a single correct Joback boiling-point value.Correct answers required functional-group decomposition, contribution retrieval, summation, unit conversion, and rounding; several models hallucinated constants.
  • General skincare scenarios: Skincare responses commonly emphasized gentle handling, barrier protection, hydration, occlusive layering, and dermatologist consultation, while some scenarios showed little consensus.Sun protection dominated pigmentary-change and anti-aging recommendations, whereas acne and excess-oil guidance consistently favored gentle cleansing.

4 Discussion and Conclusions

LLMs performed well on familiar retrieval-based cosmetic chemistry tasks and broadly sensible skincare questions, but struggled with structural reasoning, algorithmic precision, and technically specific guidance. Their fluent, confident responses may mislead consumers, so they should remain supplementary tools while developers prioritize scientific grounding and consumers seek expert verification.

  • The Disparity Between Retrieval and Reasoning: LLMs showed strong performance on familiar retrieval-based tasks but significant failures on structural reasoning and algorithmic precision.They could retrieve chemical constants with confidence but were not reliably capable of applying them through multi-step structural reasoning.
  • General Skincare Performance: LLMs gave broadly sensible guidance on general skincare topics, but responses lacked the clinical specificity required for personalized dermatological guidance.The stronger performance likely reflects the abundance of skincare-related content in their training corpora.
  • Hallucination and Consumer Safety: Plausible-sounding but factually incorrect numerical or chemical data created a risk of unwarranted consumer trust in confident answers.This concern is heightened in cosmetic and skincare contexts, where consumers may interpret authoritative responses as reliable.
  • The Role of Model Scale and Architecture: Model scale did not reliably predict performance, with smaller “flash” or “mini” models sometimes matching or outperforming larger counterparts.Training-data quality and architectural efficiency may be as consequential as raw parameter count.
  • Conclusions: LLMs are not yet reliable sources for cosmetic chemistry or precision skincare guidance despite their fluency in general advice.The gap between conversational fluency and rigorous chemical reasoning creates a superficial appearance of expertise that could mislead consumers.
  • Practical Recommendations: Developers should prioritize scientific grounding, while consumers should treat LLM-generated information as a starting point and seek expert guidance for precise chemical or safety questions.Recommended improvements include specialized chemical datasets, domain-adaptive pre-training, and clearer distinctions between evidence-based science and anecdotal content.

A Appendix: Experimental Prompts and Target Stimuli

The appendix defines benchmark prompts spanning quantitative cosmetic-chemistry inquiries and qualitative skincare scenarios. Quantitative tasks test structural, compositional, molecular-weight, and boiling-point reasoning, while qualitative tasks test immediate skincare recommendations.

  • A.1 Quantitative Inquiries: Bond Analysis asks models to state the number of sigma and pi bonds in a specified ingredient.The response must be given in one simple sentence.
  • A.1 Quantitative Inquiries: Preservative Identification asks models to list all synthetic preservatives found in a specified product.The response must be given in one sentence.
  • A.1 Quantitative Inquiries: Molecular Weight asks models to provide a specified ingredient’s molecular weight rounded to one decimal place.The response must be given in one sentence.
  • A.1 Quantitative Inquiries: Joback Calculation asks models to calculate a specified ingredient’s normal boiling point in Kelvin using the Joback method.The answer must be provided as a whole number in one sentence.
  • A.2 Qualitative Scenarios: Sequential skincare advice asks for the immediate next step after applying a scenario for sensitive, acne-prone skin.The response should be brief and limited to one sentence.
  • A.2 Qualitative Scenarios: First-step recommendations ask for the immediate first facial skincare step for treating a specified scenario.The response should be brief and limited to one sentence.

B Appendix: Correct Answers for Quantitative Inquiries · B.1 Calculating Sigma and Pi Bonds

The appendix provides manually calculated sigma- and pi-bond counts for selected cosmetic ingredients, with values independently verified through structural decomposition and cross-referenced with NCBI chemical databases.

  • B Appendix: Correct Answers for Quantitative Inquiries: The appendix includes correct answers for quantitative inquiries involving sigma- and pi-bond calculations.
  • B.1 Calculating Sigma and Pi Bonds: Table 2 presents manual calculations of sigma- and pi-bond counts for selected cosmetic ingredients.
  • B.1 Calculating Sigma and Pi Bonds: The calculations address selected cosmetic ingredients rather than an unspecified set of compounds.
  • B.1 Calculating Sigma and Pi Bonds: The reported values were independently verified through structural decomposition.
  • B.1 Calculating Sigma and Pi Bonds: The verification used structural decomposition as an explicit methodological check.
  • B.1 Calculating Sigma and Pi Bonds: The verified values were cross-referenced with NCBI chemical databases.

B.2 Identification of Synthetic Preservatives

This section presents the identification of synthetic preservative systems in selected commercial formulations. Ingredient profiles were cross-referenced against official manufacturer disclosures and the INCIDecoder database.

  • The analysis identifies synthetic preservative systems in selected commercial formulations.
  • The evaluated products were commercial formulations selected for preservative-system identification.
  • Ingredient profiles were cross-referenced against official manufacturer disclosures and the INCIDecoder database.

B.3 Calculating Molecular Weight

This section compares the molecular weights of selected cosmetic ingredients. Reported values were independently verified by stoichiometric summation and cross-referenced with PubChem.

  • Table 4 compares the molecular weights of selected cosmetic ingredients.
  • Values were independently verified through stoichiometric summation of standard atomic weights and cross-referenced against the PubChem database.

B.4 Calculating Normal Boiling Point using the Joback Method

This section estimates the normal boiling points of cosmetic ingredients using the Joback method. The calculated values use Tb = 198 + ∑∆Tb,i and are rounded to the nearest whole number.

  • Table 5 reports estimated normal boiling points for cosmetic ingredients calculated with the Joback method.
  • The Joback calculation defines boiling point as Tb = 198 + ∑∆Tb,i.
  • Calculated boiling-point values are rounded to the nearest whole number.
Loading 2608.14631v1…