Source-linked AI summary

(Whose defaults?) Is artificial intelligence reorienting archaeological methods?

Lorenzo Cardarelli, Roberto Ragno

arXiv:2609.11198v1cs.CYcs.AIcs.CLcs.HC

TL;DR

The paper examines how LLM adoption may affect archaeological methods and scientific workflows. It combines compositional modelling of computational techniques with evidence about methodological convergence, finding no demonstrated reduction in diversity but conditions consistent with a future shift toward convergence.

  • Problem

    The paper asks whether LLM adoption is narrowing archaeology’s methodological diversity and how these tools should be integrated into interpretative scientific practice.

  • Method

    The study models computational-method composition within archaeological sub-disciplines using a Bayesian hierarchical Dirichlet-Multinomial regression.

  • Results

    No individual method showed a credible change, while the aggregate post-2023 shift was smaller than the field’s existing variation and has not yet reduced methodological diversity.

  • Takeaways & Limitations

    LLMs should be integrated into complete scientific workflows rather than used as naive exploratory prompts, with analytical responsibility remaining an explicit concern.

  • Takeaways & Limitations

    The study’s corpus depends on a periodically revised source list retrieved on 29 April 2026, and its educational implications extend beyond currently demonstrated method-use changes.

Abstract

from arXiv · show

Generative AI and the practice of "vibe coding" are changing how archaeologists carry out computational research, but their effects on the discipline's range of methods is still understudied. In this paper, we evaluate whether large language models (LLMs) are narrowing the variety of methods archaeologists use. We first analysed approximately 119,000 archaeology abstracts from Scopus, covering publications from 2010 to 2025. Using a locally run LLM, we identified the computational methods reported in each abstract and organised them into 25 broad categories (L2) and 241 finer clusters (L3). A Bayesian Dirichlet-multinomial model of method composition within sub-disciplines found a small but credible shift in method use after 2023. However, this shift was smaller than the variation already present across the full study period. No individual technique showed a significant change, and overall methodological diversity increased rather than declined. We then ran a controlled experiment to see whether LLMs recommend a narrower set of methods than archaeologists have used in practice. Two different open-weight models were asked to suggest methods for 28 archaeological research problems, with prompts providing three levels of methodological guidance: novice, intermediate, and expert. Recommendation diversity was much lower than in the published literature, particularly without methodological guidance. The models also tended to favour methods that were widely used before 2023, and their recommendations more closely resembled the post-2023 literature. Taken together, these results are consistent with LLMs pushing methodological choice towards convergence, although our study cannot establish a causal effect. They raise a broader question: how can archaeology retain methodological diversity as LLMs become more involved in research?

1 Introduction

The paper examines whether LLMs narrow archaeological methodological choice, alongside the epistemological consequences of delegating computational formalisation. It combines corpus analysis with controlled method recommendations and argues for integrating LLMs into, rather than externalising, scientific workflows.

  • 1 Introduction: Vibe coding lowers barriers to custom pipelines, interfaces, and visualisations, especially for researchers without formal programming training or dedicated software teams.The authors describe this as a significant practical and potentially democratising development.
  • 1 Introduction: The paper distinguishes externalising knowledge through established tools from bypassing knowledge when LLMs determine analytical categories, variables, distances, or thresholds.A specified K-means script preserves interpretative assumptions, whereas an open classification prompt leaves them opaque.
  • 1 Introduction: Delegating code construction can create a secondary black box in which researchers produce functioning software without understanding its internal logic, undermining accountability and open-science principles.The paper links this risk to algorithmic agency and models’ tendency to favour common methods from aggregated scientific production.
  • 1 Introduction: The study analyses roughly 120,000 Scopus abstracts, models method composition with Bayesian Dirichlet-Multinomial regression, and compares LLM recommendations across standardised archaeological problems.Methods are organised into hierarchical categories, while recommendation concentration is assessed against published-literature diversity.
  • 1 Introduction: LLMs may converge archaeological method choice toward statistically common techniques, but the study tests this through both published methods and model recommendations rather than establishing causation.The mean-collapse hypothesis predicts stronger defaulting to prevalent methods when prompts provide minimal methodological guidance.

4 Results

The literature became more methodologically diverse after 2023, while LLM recommendations covered a substantially narrower range and favoured established methods. This contrast is consistent with convergence in model-guided choices, but the analyses compare different research stages and do not establish causation.

  • Literature composition: No individual method showed a credible post-2023 shift, although an aggregate change in method composition was detectable.
  • Literature composition: The effective number of L3 methods increased from 87.6 to 111.2 after 2023, indicating greater rather than lower methodological diversity in the literature.The 90% credible intervals were 84.6–90.6 before and 107.7–114.7 after 2023.
  • LLM recommendations: With neutral prompts, Qwen3 and Gemma recommended effective method sets of only 31.6 and 28.8, roughly one third of the literature’s diversity.Qwen3 was marginally more diverse than Gemma, and the gap was credible across comparisons.
  • LLM recommendations: Recommendation diversity narrowed to about 21 methods for novice prompts but rose with methodological guidance, reaching 60.7 for Qwen3 and 46.8 for Gemma under expert profiles.Even expert-profile recommendations did not match the breadth of the published literature.
  • Recommendation defaults: Both models concentrated on methods such as network analysis, modelling, simulation, and GIS, often tracking pre-2023 prevalence rather than post-2023 momentum.The negative-binomial analysis found positive associations between pre-2023 prevalence and recommendation frequency across models and profile combinations.
  • Implications and scope: The discussion links underspecified prompts and unequal auditing capacity to methodological homogenisation, while noting that the bibliometric and prompting analyses measure different research stages.The study therefore treats convergence as a potential mechanism rather than a demonstrated causal effect in the published record.

6 Conclusions and Prospects

The study finds no demonstrated decline in methodological diversity, but concentration analysis suggests conditions for convergence may already exist. It frames LLM adoption as a reconfiguration of epistemic practice requiring attention to transparency, reproducibility, and structured use.

  • 6 Conclusions and Prospects: Archaeological literature has not yet shown fewer diverse methods after LLM adoption, although concentration analysis indicates conditions for convergence may be emerging.The authors specifically link this possibility to unclear user engagement with models.
  • 6 Conclusions and Prospects: LLM-assisted code writing is presented as a reconfiguration of epistemic practice that blurs technical skill and theoretical reasoning, rather than simple tool substitution.
  • 6 Conclusions and Prospects: AI-assisted analysis is difficult to audit and cannot be fully reproducible even when prompts are disclosed because the tools are stochastic.The authors discuss structured prompting and locally hosted or open-source models as partial responses.
  • 6 Conclusions and Prospects: The paper emphasizes that computational research outputs and prompts are released, while proprietary Scopus metadata and source identifiers cannot be redistributed.The released artefacts support reproduction of downstream analyses for readers without Scopus access.

S2: Model cards and decoding settings

The supplementary material consolidates the models, decoding settings, and pipeline stages used across the machine-learning stream. The supplied passages identify the experiment-generation configuration but do not provide complete readable settings for every listed stage.

  • S2: Model cards and decoding settings: The supplementary table organizes models and decoding settings by pipeline stage, consolidating parameters distributed across the main text and supplement.
  • S2: Model cards and decoding settings: Qwen3.5-9B-Instruct is listed for experiment generation with Q8_0 quantization, context length 4,096, temperature 0.3, seed 42, and 30 maximum tokens.
  • S2: Model cards and decoding settings: The pipeline uses Qwen3.5-9B-Instruct and Gemma 4 E4B-it across its model settings.

S3: Scopus retrieval protocol

The retrieval protocol defines an archaeology corpus from Scopus sources classified under ASJC code 1204 and retrieves records year by year when necessary. The reproducible cleaned corpus contains 119,327 articles, while the source list and metadata remain subject to licensing and retrieval-date constraints.

  • S3: Scopus retrieval protocol: 546 unique Scopus Source IDs were identified using the ASJC 1204 Archaeology classification.The procedure uses Scopus Search API requests with complete views and a page size of 25.
  • S3: Scopus retrieval protocol: When a source exceeded the API ceiling, retrieval was split by publication year from 2010 through 2025, with throttling, retries, backoff, and checkpointing.
  • S3: Scopus retrieval protocol: 119,327 articles constitute the reproducible cleaned corpus used for analysis.Intermediate retrieval counts are intentionally not treated as stable because they depend on Scopus index state.
  • S3: Scopus retrieval protocol: The 8,404 articles retained at extraction step 4 produced 17,229 article–method pairs, averaging 2.05 extracted methods per article.
  • S3: Scopus retrieval protocol: The corpus may differ when regenerated later because Scopus periodically revises its source list, although the ASJC 1204 filtering procedure is deterministic.

S4: Method extraction and taxonomy construction

The method-extraction pipeline uses one local LLM call per abstract, normalizes extracted terms, and builds a fine-grained taxonomy before assigning clusters to broader categories. A conservative garbage-detection pass reduces 242 initial clusters to the 241 used in analysis.

  • S4: Method extraction and taxonomy construction: One LLM call per abstract extracts computational methods from the 119,327-article corpus using explicit-name and specificity rules.The prompt excludes software and programming languages and avoids inferring techniques not named in an abstract.
  • S4: Method extraction and taxonomy construction: The pipeline normalizes spelling, hyphenation, capitalization, and plurals before grouping equivalent method terms.The procedure uses token-sort similarity at 88/100 and retains the most frequent variant as the canonical term.
  • S4: Method extraction and taxonomy construction: The canonical vocabulary contains an over-stripping artefact in terms ending with -sis, including Principal Component Analysi.The authors state that the deformation is applied uniformly and does not split or merge otherwise distinct clusters.
  • S4: Method extraction and taxonomy construction: The taxonomy labels fine-grained clusters with short technique-family names and assigns each cluster to one broader methodological category.The assignment prompt prioritizes algorithmic or statistical paradigm over archaeological application.
  • S4: Method extraction and taxonomy construction: A conservative LLM garbage check flagged exactly one of 242 initial L3 clusters, leaving 241 clusters for analysis.The flagged cluster contained 119 canonical terms and affected 145 article–method pairs, while removal reduced the article count from 8,404 to 8,398.

S5: The full taxonomy

The taxonomy organises archaeological computational methods into 25 researcher-defined L2 categories and 241 data-driven L3 clusters, supported by definitions and representative examples.

  • Taxonomy structure: The taxonomy contains 25 fixed L2 categories and 241 corpus-derived L3 clusters, with L1 serving only as an editorial grouping layer.L2 definitions and representative methods are supplied to prompts, whereas L3 clusters are discovered from the corpus.
  • Counting conventions: Counts distinguish method mentions from distinct articles, so an article reporting multiple methods from one L2 category is counted once in article totals.The analysis corpus excludes the garbage cluster.
  • Methodological coverage: The L2 categories span statistical, spatial, temporal, imaging, machine-learning, archaeometric, environmental, simulation, network, and dimensionality-reduction approaches.Examples include Bayesian inference, spatial statistics, remote sensing, photogrammetry, supervised classification, deep learning, archaeometry, agent-based modelling, and clustering.
  • Methodological coverage: The taxonomy also covers domain-specific computational approaches for dating, morphometrics, text mining, isotope analysis, geoarchaeology, network analysis, regression, and unsupervised learning.These categories connect computational techniques to archaeological materials, environments, texts, biological data, and social relationships.

S6: LLM experiment: design and quality control

The experiment tests how methodological framing affects LLM recommendations by holding archaeological content constant while varying researcher guidance across 28 questions and three profiles.

  • Stimulus design: The experiment uses 28 thematic research questions, each formulated without computational terminology to avoid directly priming method recommendations.The questions cover archaeological topics ranging from architecture and burials to chronology, landscape, archaeobotany, trade, heritage, and rock art.
  • Prompt profiles: The three profiles represent novice, intermediate, and expert methodological guidance, with the expert profile receiving bare L2 category names.Intermediate prompts provide vague methodological descriptors, while novice prompts provide no specific computational background.
  • Experimental control: Each of 252 question–category–descriptor iterations was submitted to all three profiles, producing within-iteration comparisons that differ only in methodological framing.The experiment generated 756 calls per model, plus mapping calls for extracted L4 methods.
  • Prompt and mapping procedure: The prompts ask models to list specific computational methods with one-sentence justifications, while a separate classifier maps each generated L4 method to one L3 category.The same system framing and output-format instruction were used across profiles.

UNPARSED

The generated method strings mapped almost entirely onto the existing taxonomy, providing broad L3 coverage with only one malformed out-of-taxonomy attempt.

  • Mapping coverage: Against the 241 analysable labels, coverage was 193 labels for Qwen and 166 for Gemma because both models used the garbage cluster.The garbage cluster is excluded from the analysable taxonomy.
  • Quality control: Exactly one mapping attempt fell outside the taxonomy: Gemma produced L3-27, a malformed rendering of the three-digit code L3-027.The malformed output was attributed to a malformed L3 label rather than an otherwise new method category.

S7: The Dirichlet-Multinomial model

The Bayesian Dirichlet–Multinomial model estimates compositional method changes within L2 sub-disciplines, separating pre-existing trends from post-2023 deviations while accounting for uncertainty and implementation constraints.

  • Model specification: The model represents method counts compositionally within each L2 sub-discipline, so increasing one technique’s share requires decreases elsewhere in that composition.It includes baseline log-shares, linear year trends, post-2023 level shifts, and a Dirichlet–Multinomial concentration parameter.
  • Model interpretation: The post-LLM parameter γ measures deviation from the pre-existing linear trend, but three post-2023 years cannot distinguish an abrupt change from gradual acceleration.The primary estimand is the across-technique scale σγ rather than any individual γ.
  • Implementation details: The model uses weak sum-to-zero constraints to avoid assigning a special reference technique and tight priors for unused padding parameters to improve sampling efficiency.The count array is padded to Kmax = 20 even though sub-disciplines contain different numbers of techniques.
  • Diagnostics: Tree-depth saturation reduced sampling efficiency without biasing exploration, with bulk ESS for σγ equal to 640 under adapt_delta = 0.99.Diagnostics were reported for the whole parameter vector.
  • Posterior results: σγ is credibly above zero, but σβ exceeds σγ by a factor of 2.35, indicating that post-2023 reshuffling is smaller than the field’s historical variation.The posterior mean of ϕ is near 1,100, placing the likelihood close to multinomial and making the small shift detectable.
  • Posterior results: None of the 241 analysed L3 methods has a γ 90% credible interval excluding zero, so the aggregate signal does not localise to a named technique.Directional method lists are therefore indicative rather than individually significant.

S8: Bayesian workflow checks

The Bayesian workflow checks found that the priors covered the observed diversity range, the model reproduced all fitted sub-disciplines, and σγ was identifiable rather than prior-driven.

  • The prior predictive 5th–95th percentile range for the inverse Simpson index was [2.76, 11.57], covering the observed range [1, 12.76].
  • The analysis used posterior predictive simulations to assess calibration and prior-predictive simulations to test whether the priors covered plausible diversity values.
  • All 25 sub-disciplines fell inside the 0.05–0.95 predictive-calibration band, indicating that the model reproduced each fitted group’s compositional structure.
  • Fake-data recovery showed σγ was identified despite a wide posterior for a single sub-discipline, with aggregation across 25 sub-disciplines sharpening the full-model estimate.

S9: The concentration analysis

The concentration analysis found that LLM recommendations were substantially less diverse than the archaeological literature, while literature diversity increased rather than contracted after 2023.

  • A uniform Dirichlet prior added one pseudo-count to every L3 category, making the reported gap conservative because it inflated diversity in concentrated conditions.
  • The inverse Simpson index converted recommendation frequencies into an effective number of equally frequent methods, with lower values indicating greater concentration.
  • Gemma’s effective method counts were 28.8 overall, 20.5 for novice, 29.5 for intermediate, and 46.8 for expert prompts.
  • The literature diversified rather than contracted after 2023, whereas every LLM condition remained below both literature figures.
  • Literature posteriors had higher effective method counts than every Qwen3 and Gemma profile, with no overlap; within each model, diversity increased from novice to expert prompts.

S10: What drives LLM recommendations

The recommendation-count analysis separated effects of pre-2023 prevalence from post-2023 movement and found that models preferentially recommended techniques prominent in the earlier literature, especially under novice prompts.

  • The negative-binomial model included both pre-2023 prevalence and post-2023 movement so training-distribution mimicry was not conflated with recent disciplinary change.
  • Pre-2023 prevalence was credibly positive in all eight fits, with posterior probability 1.00 of exceeding zero, and the association weakened as methodological guidance increased.
  • The prevalence gradient was steeper for Qwen3, declining from 1.039 to 0.419, than for Gemma, declining from 0.675 to 0.511 across prompt profiles.
  • The post-2023 movement coefficient was indistinguishable from zero, but this null result was expected to be uninformative because the predictor depended on noisy, individually non-credible γ estimates.

S11: Robustness and sensitivity

Sensitivity analyses supported the main qualitative pattern: the field changed internally without measurable sub-discipline homogenisation, and LLM recommendations resembled post-2023 literature more than pre-2023 literature.

  • Changing the breakpoint from 2023 to 2022 left qualitative conclusions unchanged, including no credible change for any individual technique.
  • Restricting the analysis to 33 well-represented techniques produced another null overall coefficient, −0.138, with P(β > 0) = 0.438.
  • The distributional test found LLM recommendations credibly closer to post-2023 than pre-2023 literature, with weakest methodological guidance producing the strongest alignment.
  • The distributional result establishes similarity rather than causation because models may reflect existing training-data trends, influence adoption, or both.
  • All 25 sub-discipline γ intervals straddled zero, while between-sub-discipline variation exceeded post-2023 change by an order of magnitude.

S12: Reproducibility

The supplementary materials document the corpus flow, modelling inputs, software, and scripts needed to reproduce the analyses. Reproduction from raw Scopus exports remains constrained by proprietary access, while later analyses can use deposited artefacts.

  • Reproducibility workflow: The reproducibility workflow carries the corpus through an aggregated Stan input array, with every figure re-derived from deposited artefacts by a verification script.The deposited array is produced after an additional modelling-stream filter.
  • Corpus construction: The analysis corpus is filtered before modelling, with 2026 partial-year records and pre-2010 articles removed, leaving 14,138 observations split into pre-2023 and post-2023 periods.The retained observations comprise 8,442 pre-2023 and 5,696 post-2023 records.
  • Taxonomy and analysis inputs: The taxonomy distinguishes 242 generated L3 clusters from the 241 clusters used in analysis, while recommendation coverage uses 242 as its denominator.The extra cluster is retained for coverage reporting because both models mapped items into it.
  • Software and scripts: The supplementary scripts cover data preparation, Dirichlet-multinomial modelling, posterior extraction, technique-level effects, workflow checks, robustness, concentration, and literature-versus-LLM regression.The listed scripts are divided between bibliometric and experiment streams.
  • Access constraints: Raw Scopus metadata and source identifiers are not redistributed, so reproducing the initial bibliometric stream from scratch requires Scopus access; downstream analyses can use deposited files.The deposited materials include the aggregated count array, taxonomy, stimuli, and model recommendations.

List of supplementary tables

The supplementary table list maps the study’s supporting materials to corpus construction, taxonomy, research questions, and modelling outputs. It also identifies the software and corpus-flow tables used for reproducibility.

  • Machine-learning stream: The listed supplementary tables cover computational environments, model settings, Scopus metadata, corpus construction, clustering, taxonomy definitions, research questions, and recommendation materials.The list includes tables from S1 through S6.

List of supplementary figures

The supplementary figure list highlights posterior method effects and distributions of effective methodological diversity across conditions.

  • Method effects: The listed supplementary figures show posterior γ estimates for the 20 L3 methods with the largest absolute effects, including 90% credible intervals.These figures are identified as S7.1 and S7.6.
  • Methodological diversity: The figure list also includes posterior distributions of the effective number of L3 methods by condition.These figures are identified as S9.1 and S9.2.
Loading 2609.11198v1…