Source-linked AI summary

When Literature Data Mislead Artificial Intelligence in Materials Discovery

Qian Wang, Ying Li, Ryuhei Sato, Hidemi Kato, Shin-ichi Orimo, Hao Li, Eric Jianfeng Cheng

arXiv:2609.01621v1cs.IRcond-mat.mtrl-scics.CEcs.LG

TL;DR

Literature-derived solid-electrolyte conductivity data can contain ambiguous reporting and unit errors that are difficult to detect yet matter for AI reuse. The paper traces reported values against plotted and converted values, finding a cross-database example with a 100-fold conductivity error.

  • Problem

    Ambiguous reporting and unit inconsistencies in solid-electrolyte literature can remain embedded in datasets despite being difficult for domain experts to detect.

  • Method

    The paper traces conductivity values by comparing reported values with values obtained through replotting and unit conversion.

  • Results

    A cross-database example shows that ambiguous conductivity reporting can produce a 100-fold error.

  • Takeaways & Limitations

    Accurate, traceable data are presented as prerequisites for reproducible science and trustworthy AI-assisted discovery.

  • Takeaways & Limitations

    The analysis identifies representative and reproducible inconsistency types rather than providing an exhaustive statistical survey of the entire solid-electrolyte literature.

Abstract

from arXiv · show

Artificial intelligence (AI) increasingly treats scientific literature as a data source for building databases, training predictive models, and guiding discovery. Yet literature-derived datasets often assume that reported experimental values are internally consistent and directly reusable. Here, we analyze this assumption using solid electrolyte (SE) conductivity data as a representative materials-science case. By tracing values from source articles to curated datasets, we identify recurrent text-figure mismatches, ambiguous axis annotations, unit inconsistencies, and missing measurement context. These discrepancies are often numerically plausible and therefore difficult to detect through routine preprocessing, but they can propagate as structured label noise during database construction and machine-learning reuse. A cross-database example shows how ambiguous reporting can create a 100-fold conductivity error. Our analysis reframes data accuracy as an infrastructure requirement for artificial-intelligence-driven discovery and motivates traceable reporting, curation, and validation practices for reusable scientific data. Keywords: AI for science; Data reliability; Scientific databases; Structured label noise; Literature-derived data; Materials informatics; Solid electrolytes

Results 97

Curation of solid-electrolyte conductivity literature reveals text–figure mismatches, annotation and unit ambiguities, and incomplete measurement context that can produce systematic database-label uncertainty. These numerically plausible errors can propagate across databases, underscoring the need for traceable reporting and interpretive curation before AI reuse.

  • Reporting problems: Four recurring reporting problems are text–figure inconsistency, annotation ambiguity, unit inconsistency, and incomplete measurement context.These issues particularly affect ionic conductivity and other Arrhenius-derived quantities, including activation energy.
  • Text–figure inconsistency: 7.2×10–4 S cm–1 at 25 °C in the text corresponds to a plotted point closer to 30 °C.Repeated text–figure mismatches introduce systematic uncertainty into database labels rather than remaining local presentation errors.
  • Traceable curation: Curation requires reconstructing numerical meaning before standardization, because reusable data need traceability for the reported value, its physical meaning, and the curation decision.Ambiguous entries can remain numerically plausible and embedded in datasets, while cross-database propagation limits data-quality guarantees from database construction alone.

Methods 255

The study curated experimental solid-electrolyte conductivity data from literature and databases, independently re-extracted values, and classified reporting ambiguities through cross-source validation. The analysis was representative rather than exhaustive and relied partly on expert judgment because graphical interpretation depends on resolution and reporting clarity.

  • Dataset construction: The curated dataset comprised 3,814 experimental solid electrolytes with more than 27,600 conductivity entries, including inorganic, polymer, and gel electrolytes.The dataset included 2,900 inorganic, 528 polymer, and 386 gel solid electrolytes.
  • Independent reanalysis: Experimental conductivity values were independently re-extracted from figures and tables by reconstructing conductivity–temperature relationships and checking axes, scales, annotations, units, and conversions.Temperature values were converted between °C and K and mapped to Arrhenius coordinates when necessary; plausible plotted quantities were considered when axis definitions were ambiguous.
  • Ambiguity classification: Reporting issues were classified as text–figure inconsistency, annotation ambiguity, unit inconsistency, or incomplete measurement context.Cases with multiple issue types received all applicable classifications rather than being forced into one category.
  • Case identification and statistical analysis: 25 issue occurrences arose from 16 inorganic source cases, while gel-polymer statistics were calculated at the paper level from 83 screened publications.Inorganic statistics used issue occurrence because one source case could contain multiple issue types.
  • Scope and limitations: The analysis targeted representative and reproducible ambiguity types rather than an exhaustive survey, and graphical interpretation could depend on figure resolution, reporting clarity, and expert judgment.Curated databases were also treated as potentially reflecting prior literature inconsistencies and requiring expert validation.

Supplementary Information 416

The supplementary information catalogs solid-electrolyte papers affected by annotation, unit, text–figure, and measurement-context problems. It also summarizes DigBat’s temporal and categorical coverage of inorganic, polymer, gel, and computational solid-electrolyte data, including unique DOI counts by category.

  • Issue catalog: The supplementary tables organize gel and inorganic solid-electrolyte papers by case ID, DOI, and issue category.The listed issue categories include annotation ambiguity, unit inconsistency, text–figure inconsistency, and incomplete measurement context.
  • Issue catalog: The inorganic cases repeatedly combine annotation ambiguity or text–figure inconsistency with unit inconsistency.Examples include cases I3, I5, I7, I8, I10, and I11, while I13 additionally involves incomplete measurement context.
  • Issue catalog: The catalog also records isolated unit inconsistencies and annotation ambiguities across additional inorganic solid-electrolyte cases.Cases I14–I16 are listed with unit inconsistency, unit inconsistency, and annotation ambiguity plus unit inconsistency, respectively.
  • DigBat database overview: DigBat statistics track the temporal evolution of recorded materials across inorganic, solid polymer, gel polymer, and computational data.The accompanying inset summarizes unique DOIs associated with each category, and all statistics are derived from DigBat.
Loading 2609.01621v1…