Source-linked AI summary
Process-Constituted Intelligence: A Shared Criterion for Humans and Machines
Michael J. Richardson, Ayeh Alhasan, Cassandra Crone, M. Paula Diaz Monfort, Patrick Nalepka, Mark Dras, Rachel W. Kallen, David M. Kaplan
TL;DR
GenAI reproduces traces of human cognition while largely omitting the processes that generate them. This paper defines seven-feature strong equivalence, proposes process-based designs and audits, and concludes that current systems often produce process-shaped output without instantiating the depicted process.
Problem
GenAI is trained on recorded cognitive traces while iterative, dialogical, and tacit processes behind them remain largely absent, creating a gap between output resemblance and cognitive process.
Method
The paper defines strong equivalence across seven process features and proposes process-engineering commitments plus substrate-neutral audits for comparing human and machine activity.
Results
Current GenAI produces process-shaped output whose depicted process its architecture does not instantiate, and no existing scaffold covers all seven features.
Takeaways & Limitations
Process-focused designs and audits shift evaluation from output appearance toward preserving and measuring the generative activity underlying judgment and creativity.
Takeaways & Limitations
No existing scaffold covers all seven process features, leaving people to supply the features current architectures lack.
Abstract
from arXiv · showhide
Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself. Generative AI (GenAI) is trained on \textit{traces} (textual and visual residues of human cognitive processes), reproducing samples from a distribution of those traces. Its outputs resemble reasoning, problem-solving, and creativity, yet the activity that produces such outputs in humans remains largely absent. Current GenAI is, therefore, weakly equivalent to the cognition it imitates, matching outputs while process stays absent or opaque. The cognitive sciences have long distinguished between weak and strong equivalence. Here, we define \textit{strong} equivalence across seven process features, assessable against human and machine cognition. Our process-based account addresses a symmetric risk: GenAI tools that outsource a person's generative processes may leave critical capacities unbuilt. We specify design principles for GenAI that instantiate more process and preserve rather than erode human judgment and creativity, and outline process audits that make strong equivalence testable.
1 The output-delivery pattern
Current GenAI primarily delivers outputs that can resemble human reasoning, problem-solving, and creativity while leaving the processes that generate them absent, unreliable, or opaque. The paper therefore proposes a process-level criterion for strong equivalence, applies it to human and machine cognition, and uses it to motivate process-preserving designs and audits.
- The output-delivery pattern: GenAI primarily follows an output-delivery pattern in which a request produces an answer, paragraph, image, or code block resembling competent human work.The underlying human activity—including drafts, abandoned attempts, pauses, and dialogue—is largely absent from the delivered output.
- The output-delivery pattern: Models trained on traces can exhibit weak equivalence by matching human input-output behavior without reproducing the computational processes that generated it.Reasoning-shaped text and reasoning traces may be unreliable or partial windows onto the processes driving answers.
- The output-delivery pattern: Agentic architectures and iterative reasoning methods move process toward the center but remain weak and partial without substrate-neutral principles for measuring or auditing cognitive processes.The paper identifies a need for process criteria that can assess cognition across machine systems and brains without requiring a shared algorithm.
- The output-delivery pattern: Outsourcing generative work to GenAI may leave the human capacity that work normally requires neither instantiated nor refined, motivating process-preserving design principles and matched-task process audits.The proposed framework reads current GenAI through a process criterion, addresses AI-assisted human work, and specifies audits across both substrates.
2 Intelligence as process
Intelligence is constituted by the iterative activity of producing, testing, and revising rather than by products or stored competence inferred from results. These capacities are acquired through structured practice, and process-constituted intelligence comprises seven observable features expressed across substrates as one iterative cycle.
- 2.1 The process and the trace: Intelligence refers to the activity through which capacities are exercised, not merely to a stored competence read off its results.
- 2.1 The process and the trace: Structured, iterative practice builds capacity through attempting, failing, and revising; exposure to finished products cannot substitute for having performed that activity.
- 2.2 Process features: Process-constituted intelligence has seven core features, each observable in solver behavior and instantiated differently across material substrates.
- 2.2 Process features: Generative trial and revision treats producing many attempts and refining through failures as cognitive work, including developing judgment about promising moves and worthwhile problems.
- 2.2 Process features: Social and dialogical accountability makes cognition answerable to interlocutors and to the normative demand for getting things right.
- 2.3 One cycle, many substrates: The features form one iterative cycle: a solver generates moves, meets resistance, reframes, revises, and repeats across many turns.
3 Weak, strong, and hard equivalence
The section distinguishes matching cognitive outputs from instantiating the processes that produce them, extending weak–strong equivalence into three grades for contemporary AI. It illustrates the distinction with Centaur, whose human-like predictions remain only weakly equivalent because its process differs from and is less robust than human cognition.
- 3.1 The distinction: Weak equivalence means computing the same input–output function, whereas strong equivalence requires the same algorithm in the same functional architecture.The distinction separates behavioral resemblance from process and architectural identity.
- 3.1 The distinction: The Turing test certifies only weak equivalence because passing systems match behavior without establishing that they share the underlying process.Behavioral success alone does not determine how the behavior was produced.
- 3.2 Three grades, not two: Contemporary AI analysis supplements Pylyshyn’s binary distinction with an intermediate grade between weak and hard equivalence.The original strong equivalence is relabeled hard equivalence and remains algorithmic-architectural.
- 3.3 A current illustration: Centaur was fine-tuned on over ten million human choices across hundreds of experiments and predicted held-out human behavior better than bespoke models, yet remained only weakly equivalent.Its predictive power does not show that it arrives at outputs as people do.
- 3.3 A current illustration: Centaur’s behavioral match degraded when small wording changes shifted meaning, while human respondents tracked those changes.This breakdown exemplifies the manipulations targeted by a process-level account.
4 The machine case: a weak-equivalence engine
Current GenAI is trained on human cognitive traces that largely omit the iterative activity producing them, creating a process gap that existing reasoning scaffolds only partially address. Strong-equivalent design therefore requires engineering generative activity into architectures rather than merely optimizing outputs.
- The machine case: a weak-equivalence engine: GenAI training corpora contain finished textual and visual traces while largely omitting iterative trials, abandoned approaches, dialogical pushback, and tacit shaping.The passage identifies the structure of the training signal—not simply insufficient data—as the root of the process gap.
- The machine case: a weak-equivalence engine: Contemporary scaffolds add intermediate computation through branching, critique-and-revision, code execution, and multi-agent challenge, but whether they instantiate process remains unresolved.These methods allocate inference-time compute to deliberation before answering, without establishing full process equivalence.
- The machine case: a weak-equivalence engine: Faithfulness depends on task demands and training: reasoning traces become load-bearing when tasks require serial computation and training rewards verbalization, but most deployed regimes do not compel this.Under those common conditions, GenAI produces process-shaped output whose depicted process does not necessarily perform the answer-producing work.
- The machine case: a weak-equivalence engine: No existing scaffold instantiates all seven process features, leaving humans to supply missing capacities such as worthwhile-question framing and engagement with uncertainty.Branching search, rejection sampling, and self-refinement support generative trial and revision, while framing still arrives from the prompt.
- The machine case: a weak-equivalence engine: Closing the gap is framed as process engineering: architectures should build generative activity into their operation, including dialogue that challenges problem framing and rewards substantive changes of mind.The proposed dialogical-revision architecture targets dialogical accountability and value-laden framing rather than longer chains, agreement, or voting.
5 The human case: process on the human side
When tools perform the generative activity through which human capacities form, they may produce short-term gains while eroding unaided cognition. The relevant boundary is whether assistance substitutes for or preserves the person’s generative process, a standard that guides design and motivates comparative evaluation.
- 5 The human case: process on the human side: Offloading a cognitive process can remove the activity through which a capacity forms, pushing human cognition toward the machine’s weak equivalence.What people do with tools changes what they can do without them.
- 5.2 The evidence, and its boundary: Early studies show short-term dependence patterns: unrestricted GenAI improved performance while available but worsened unaided outcomes, while developers worked more slowly despite believing they were faster.The evidence supports the framework’s prediction but does not yet establish long-term causal formative effects.
- 5.2 The evidence, and its boundary: The outcome depends on whether AI performs the generative step: answer-producing assistance harmed learning, whereas hint-based assistance that withheld answers preserved it.The evidence does not support a blanket ban on AI assistance.
- 5.3 The formative stakes: Erosive offloading threatens the slow formation of judgment, recognition of ill-posed problems, taste, and practical wisdom through difficult, uncertain, dialogical activity.These capacities are formed below articulation and are especially tempting to outsource to capable assistants.
- Design implications: Process-preserving design keeps generative activity with the person, and the field needs comparative studies testing it against process-substituting assistance on unaided capacities.Questions and counterexamples extend the process, whereas answers and ambiguity-resolving responses compress or dissolve it.
6 A shared measurement via process audits
Process audits make strong equivalence testable by measuring seven features at a resolution shared across human and machine solvers. They connect behavioral traces to empirical tests of whether systems improve process rather than only output.
- Shared audit design: A process audit is a task-and-rubric protocol that scores human or machine behavioral traces against seven features using substrate-specific anchors.Defining features more coarsely than algorithms enables the same audit to be administered across substrates.
- Shared audit design: Ill-posed reasoning, mid-trace perturbation, and dialogical-accountability probes produce behavioral signatures that can be scored on human transcripts and machine traces.The probes assess ill-posedness detection, response to new information, and engagement with a challenge’s substance rather than chain length.
- Empirical validation: An architecture implementing a feature should outperform the replaced scaffold on that feature’s audit after controlling for output quality.This makes process gains empirically distinguishable from improvements in output alone.
- Empirical validation: Higher benchmark scores without higher audit scores indicate improved output rather than improved process.On the human side, audits provide the missing dependent measure for comparative studies.
7 Conclusion
The conclusion defines intelligence by process rather than output, distinguishing weak equivalence from strong equivalence. It applies this criterion to machine architectures, human-facing tools, and process audits while contrasting process engineering with output conditioning.
- Process criterion: Intelligence is constituted in process, so matching cognitive outputs does not establish strong equivalence with cognition-constituting activity.Current GenAI is trained on traces of human cognition and is built primarily to match outputs, while process remains largely unaddressed.
- Design consequences: The same seven-feature criterion specifies both machine architectures that instantiate more process and human tools that preserve generative, uncertain, dialogical, and formative activity.The contrast is between process-preserving assistance and process-substituting assistance.
- Process audits: Process audits make the standard measurable for human and machine solvers by testing whether architectures designed around the features score higher without cosmetic compliance.The audits use probes targeting the features and score behavioral signatures on matched tasks.
- Process engineering: Process engineering targets generative activity itself, whereas prompt and context engineering optimize how systems are conditioned to improve returned outputs.The distinction matters if general capacities are constituted in process rather than read off a distribution of outputs.
- Significance: Because systems producing intelligence-like outputs are now built and deployed at scale, the difference between output and process has become consequential.The conclusion frames this as a newly urgent version of longstanding questions about what intelligence is and what it is for.
Supplementary Information
The supplementary information operationalizes the seven features of process-constituted intelligence by defining each feature for human and AI/machine cases and specifying observable diagnostic signatures. These signatures connect the feature framework to process audits that distinguish genuine presence from imitation.
- Seven process features: Supplementary Table 1 expands the seven features with a working definition, human and AI/machine instantiations, and a diagnostic signature for each.The diagnostic signature specifies what an auditor would observe to judge whether a feature is present rather than absent.
- Process audits: The diagnostic column provides an operational bridge between the seven features and the process audits described in Section 6.It maps observable evidence to audit judgments about whether a feature is genuinely present or merely imitated.
- Limitations: For current systems, the supplement indicates that some process capacities remain largely absent and should be assessed at the human–AI system level.The provided passages frame this as a limitation requiring evaluation of the combined human–AI system.