Paper deep dive
Process-Constituted Intelligence: A Shared Criterion for Humans and Machines
Michael J. Richardson, Ayeh Alhasan, Cassandra Crone, M. Paula Diaz Monfort, Patrick Nalepka, Mark Dras, Rachel W. Kallen, David M. Kaplan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/23/2026, 1:42:12 AM
Summary
The paper argues that intelligence is constituted by iterative cognitive processes rather than just outputs. It distinguishes between weak equivalence (matching outputs) and strong equivalence (matching underlying processes). The authors propose seven process features (generative trial, temporal extension, engagement with uncertainty, feedback with medium, value-laden framing, social accountability, and formative dimension) as a criterion for strong equivalence. They critique current Generative AI (GenAI) for being trained on 'traces' of human cognition, resulting in weak equivalence, and propose design principles and audits to achieve stronger process-based equivalence.
Entities (8)
Relation Signals (6)
Strong Equivalence â isdefinedby â Seven Process Features
confidence 96% ¡ Here, we define strong equivalence across seven process features, assessable against human and machine cognition.
Generative AI â exhibits â Weak Equivalence
confidence 95% ¡ Current GenAI is, therefore, weakly equivalent to the cognition it imitates, matching outputs while process stays absent or opaque.
Generative AI â istrainedon â Traces
confidence 94% ¡ Generative AI (GenAI) is trained on traces (textual and visual residues of human cognitive processes)
Centaur â exhibits â Weak Equivalence
confidence 92% ¡ Although its predictive power is impressive, the system is still only weakly equivalent.
Pylyshyn â proposed â Weak Equivalence
confidence 90% ¡ We then locate a strong equivalence between Pylyshynâs weak pole (matching outputs) and hard pole (matching algorithm and architecture)
Chain-of-Thought â isatypeof â Reasoning Scaffold
confidence 88% ¡ reasoning models often produce 'reasoning-shaped' text... chain-of-thought, ReAct, tree-of-thoughts
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself. Generative AI (GenAI) is trained on \textit{traces} (textual and visual residues of human cognitive processes), reproducing samples from a distribution of those traces. Its outputs resemble reasoning, problem-solving, and creativity, yet the activity that produces such outputs in humans remains largely absent. Current GenAI is, therefore, weakly equivalent to the cognition it imitates, matching outputs while process stays absent or opaque. The cognitive sciences have long distinguished between weak and strong equivalence. Here, we define \textit{strong} equivalence across seven process features, assessable against human and machine cognition. Our process-based account addresses a symmetric risk: GenAI tools that outsource a person's generative processes may leave critical capacities unbuilt. We specify design principles for GenAI that instantiate more process and preserve rather than erode human judgment and creativity, and outline process audits that make strong equivalence testable.
Tags
Links
- Source: https://arxiv.org/abs/2608.16213v1
- Canonical: https://arxiv.org/abs/2608.16213v1
Trouble viewing inline? Open PDF directly â
Full Text
60,535 characters extracted from source content.
Expand or collapse full text
Process-Constituted Intelligence: A Shared Criterion for Humans and Machines Michael J. Richardson1,2,4, Ayeh Alhasan1,2, Cassandra Crone1,2, M. Paula Diaz Monfort1,5, Patrick Nalepka1,2,3, Mark Dras2,4,6, Rachel W. Kallen1,2, and David M. Kaplan1,2,3 1School of Psychological Sciences, Macquarie University, Sydney, NSW 2109, Australia; 2Performance and Expertise Research Centre, Macquarie University; 3Minds and Intelligences Research Centre, Macquarie University; 4Frontier AI Research Centre, Macquarie University; 5Scuola Superiore Meridionale, Napoli, Italy; 6School of Computing, Macquarie University Correspondence: Michael J. Richardson (michael.j.richardson@mq.edu.au) Abstract Intelligence is constituted by process (iterative activity through which output emerges), not in the output itself. Generative AI (GenAI) is trained on traces (textual and visual residues of human cognitive processes), reproducing samples from a distribution of those traces. Its outputs resemble reasoning, problem-solving, and creativity, yet the activity that produces such outputs in humans remains largely absent. Current GenAI is, therefore, weakly equivalent to the cognition it imitates, matching outputs while process stays absent or opaque. The cognitive sciences have long distinguished between weak and strong equivalence. Here, we define strong equivalence across seven process features, assessable against human and machine cognition. Our process-based account addresses a symmetric risk: GenAI tools that outsource a personâs generative processes may leave critical capacities unbuilt. We specify design principles for GenAI that instantiate more process and preserve rather than erode human judgment and creativity, and outline process audits that make strong equivalence testable. 1 The output-delivery pattern â A preliminary and abridged version of some of these arguments appears in the proceedings of PAAMS 2026 50. GenAI is primarily deployed using one pattern: output-delivery. A request goes in, an output comes out (a paragraph, image, block of code, an answer). These outputs increasingly resemble the products of human reasoning, problem-solving, and creative work. For many tasks, these outputs are hard to distinguish from what a competent or expert human would produce. Yet, the activity that generates these outputs in us (the early drafts and abandoned attempts, the breaks taken at an impasse, the dialogue with collaborators and the manipulation of the material itself, the slow accrual of a sense for which moves are promising and which are dead ends) is manifestly not what the system does when it answers. A scientific breakthrough, mathematical proof, or piece of skilled craft are not simply the outputs that survive at the end of a given process, but are also constituted by the extended activity that produced them 52; 45; 23. The output itself is only a preserved record âa traceâof that activity. Intelligence, on the view we develop here, is constituted by the process and not by the trace, a distinction with important consequences for systems trained on traces and for the tools that increasingly mediate human processes through which critical thinking, judgment, and knowledge are formed. The cognitive sciences already have terminology for this distinction, differentiating between weak and strong equivalence 47. Generative models are trained on traces and reproduce samples from a distribution of those traces 6. To the extent that these outputs resemble those generated by humans performing the same task, such models exhibit weak equivalence: they reproduce the same input-output behavior without necessarily reproducing the computational processes that generated it. Strong equivalence, by contrast, requires not only matching input-output behavior but also implementing the same underlying processes that give rise to that behavior. For instance, reasoning models often produce "reasoning-shaped" text that does not track the computation driving their answers 63; 32; 10. Reasoning traces are, at best, unreliable and partial windows onto the process that generated the output 29. They do not exhibit strong equivalence. Although a move toward agentic architectures and more elaborate âreasoningâ frameworks (chain-of-thought, ReAct, tree-of-thoughts, multi-agent debate, iterative self-refinement, dedicated inference-time reasoning) 65; 68; 69; 16; 41; 44; 22 is a positive step toward putting process in central focus and moving beyond the output trace approach, frontier AI systems will remain relatively weak and partial without a clear set of substrate-neutral principles for measuring or auditing cognitive processes. The same account that finds process missing from current GenAI holds that human cognitive capacities are themselves constituted through process. A critical implication of this view is that when a human uses a GenAI tool to perform the generative work, the capacity that work would have normally required is neither instantiated nor refined in the person. Early evidence already supports this. Students given an unrestricted generative tool perform better with it and worse without it than peers who never used one 4, as the bulk of cognitive effort shifts from generation toward verification 35. Machine intelligence that is only weakly equivalent to what it imitates pushes human users strikingly towards that same weak equivalence as their own normal cognitive processes are supplanted or bypassed. These are not two separate problems, but rather one problem expressed across two substrates, for which no constructive evaluative framework yet exists. Here, we take several steps to address this issue. We do not attempt to settle what intelligence is in general (a question older than the field itself and not one we could resolve here), but aim more narrowly to specify a criterion, pitched at the level of process, that can be applied comparatively across human and machine cognition. First, we characterize process-constituted intelligence in terms of seven core features, identifiable across natural and artificial systems, yet differently realized in each. We then locate a strong equivalence between Pylyshynâs weak pole (matching outputs) and hard pole (matching algorithm and architecture) 47, at the resolution of those features, that a machine model or system (e.g., transformer) and a brain can meet without sharing an algorithm. Finally, we read current GenAI through that criterion to expose its process gap and the architectures that would close it, turn the same criterion on AI-assisted human work to derive process-preserving design principles, and propose process audits that test the framework on matched tasks across both substrates 50. 2 Intelligence as process 2.1 The process and the trace. That intelligence lies in the act of producing and not just in the product is a settled idea in the cognitive sciences, philosophy of mind, and the arts and design traditions. Throughout, we use intelligence and cognition interchangeably, referring to the activity through which such capacities are exercised rather than to a stored competence read off its results. A scientist works for months on a single question, discarding the models they were trained to expect would work before the data forced a new one. A painter works across many studies, and a mathematician returns to the same problem across years 52; 45; 23. Our claim is that intelligence is constituted just as much by the discarded studies, dead ends, and slow accrual of judgment, none of which survive in the final output the activity leaves behind, as those final "intelligent" outputs. This is not merely a stipulation about where intelligence is located, but a claim about how the capacities themselves are acquired. One does not learn to paint by studying finished canvases, nor to prove theorems by reading completed proofs. The capacity is built through structured, iterative practice (attempting, failing, and revising), and exposure to the products of that activity is no substitute for having done it 18; 33. A system, or a person, given only the traces of that activity inherits its surface and not the competence that produced it. What these traditions leave open is how to make the descriptions of these disparate processes precise enough to compare across cases. This is the task we take up here. 2.2 Seven candidate features of process. We characterize process-constituted intelligence in terms of seven core features, each of which is observable in the behavior of a solver and each instantiated differently across material substrates (Table 1; full definitions and a diagnostic signature for each feature are given in Supplementary Table 1). Generative trial and revision. Cognition proceeds by producing many attempts and refining through failures. An inventorâs failed prototypes, a mathematicianâs abandoned approaches, a painterâs discarded studies are not preliminary to the work but, rather, are the substance of it 23. Temporal extension. What matters is not that cognition takes time (all behavior does) but that it returns across occasions to the same material, carrying state forward so that each pass reworks the last, from a scientistâs months on a single problem to the years of deliberate practice that build a skill 18. Engagement with uncertainty. A solver sits in not-knowing, recognizes when a problem is ill-posed, and refuses premature closure rather than resolving to a confident answer the situation does not warrant 14. Feedback with the medium. The material (a proof, canvas, instrument, dataset) resists and redirects the activity, so the work responds to the medium rather than executing a pre-formed plan 52; 26. Value-laden framing. What counts as a promising move or worthwhile problem is constituted by the practitionerâs developing judgment and is itself cognitive work, not a parameter fixed in advance 52. Social and dialogical accountability. Cognition is constituted in part by being answerable, both to others (reasoning with and against interlocutors and having oneâs framings challenged) and, more broadly, for getting things right, a normative dimension some philosophers take to be intrinsic to intelligence 38; 2; 58; 25. A formative dimension. Much of the process operates below explicit articulation, is built up through embodied practice within a community, and simultaneously shapes who the practitioner is becoming 46; 33; 40. 2.3 One cycle, many substrates. These features are not a checklist of independent items but facets of a single iterative cycle (Fig. 1A). A solver generates candidate moves, encounters resistance from the medium or interlocutors, reframes when that resistance reveals the framing to be inadequate, revises, and repeats across many turns. This cycle is recognizable wherever cognition is constituted in ongoing activity as it runs, rather than retrieved from a store of prior results, and it is not uniquely human. Honeybee swarms select nest sites through distributed scouting, advertisement, and cross-inhibition that weighs alternatives and commits only as evidence accumulates 54. Ant colonies allocate labor and solve routing problems through local interactions that no individual represents 21. An acellular slime mold builds efficient transport networks by reinforcing productive paths and pruning others, externalizing a form of memory into the medium it moves through 43; 61; 49. These are not metaphors for cognition but instances of the same generateâencounterâreframeârevise structure realized in biological media 12; 36; 39. The same loop appears in artificial learning systems: reinforcement learning improves a policy through action, feedback, and error resolution, and policy-gradient methods reproduce the patterns of human practice-based skill learning 24; what such systems track is the change in performance across practiceâlearning rather than any single output 34. The claim is therefore not that all intelligence is human-like, but that the process is recognizable across biological, human, and machine cognition. Although each substrate instantiates it differently, this allows the same features to serve as a common standard. Fig. 1: Process, and the grades of equivalence. (A) Intelligence as process. A solver generates, encounters resistance from the medium or from interlocutors, reframes when that resistance exposes an inadequate framing, and revises. One iterative cycle whose facets are the seven features (1â7), recognizable across biological, human, and machine cognition while each substrate realizes it differently (Table 1). (B) Three grades of equivalence between a system and the cognition it models 47. Weak, the two match on outputs, achievable but silent on how either system got there. Strong (ours), they match on process (the seven features), realized differently in each substrate, the balance the rest of the paper holds systems to. Hard, they match on algorithm and architecture, Pylyshynâs original strong equivalence, which a transformer and a brain neither share nor need. Strong equivalence sits between weak (too little) and hard (too much). Table 1: Seven features of process-constituted intelligence, with biological, human, and machine instantiations. Feature Biological cognition Human cognition Machine cognition Generative trial and revision Scout bees sampling alternative nest sites; colony-level path exploration A mathematicianâs abandoned approaches; an artistâs many studies Multi-agent rollouts; tree-of-thought branching; rejection sampling with critique Temporal extension Trail reinforcement and foraging histories accrued over time Years of skill-building; a scientistâs years on one problem Persistent agentic trajectories with state across extended interactions Engagement with uncertainty Quorum thresholds that delay commitment until evidence accumulates Recognizing ill-posed problems; refusing premature closure Calibrated probabilities; explicit ambiguity flagging; abstention Feedback with the medium Stigmergic update of the environment (pheromone, network reinforcement) Painter conversing with the canvas; mathematician with the proof Tool-using agents that update framing on environmental return Value-laden framing Selection pressures defining a âgoodâ site or route Expert framing of what counts as a worthwhile problem Reward-shaping toward problem-reframing rather than solution-pursuit Social and dialogical accountability Cross-inhibition and interaction among scouts yielding a collective decision Chavruta; Socratic dialectic; peer-reviewed inquiry Multi-agent adversarial revision; inter-agent challenge Formative dimension Colony-level adaptation; developmental plasticity Becoming the kind of scientist or artist who can do this work Beyond single-system scope; appears via humanâAI co-formation (S5) 3 Weak, strong, and hard equivalence 3.1 The distinction. The distinction between resembling a cognitive process and instantiating it has long been established. Four decades ago, Pylyshyn separated weak from strong equivalence between an information-processing model and the cognition it represents 47. Two systems are weakly equivalent when they compute the same inputâoutput function and strongly equivalent when they compute it using the same algorithm in the same functional architecture (i.e., the same underlying computational organization, not merely the same inputâoutput mapping). The distinction also tracks Marrâs levels of analysis 42, with strong equivalence demanding correspondence at the algorithmic level rather than mere agreement at the computational level about what function is computed. As Marr and others have noted, matching inputâoutput behavior underdetermines the algorithm that produces it, just as a shared algorithm underdetermines its physical implementation 42. A system can reproduce the outputs of cognitive capacities while realizing different architectures entirely 20. Likewise, the Turing test certifies only weak equivalence, since systems that pass it have matched the behavior but nothing follows about the process 13. 3.2 Three grades, not two. In the context of contemporary AI, Pylyshynâs binary distinction (which was appropriate in the context of psychological explanations) is more useful if supplemented by an intermediate form of equivalence that lies between weak and hard (Fig. 1B). Pylyshynâs original strong equivalence, which we relabel hard equivalence (i.e., algorithmic-architectural), requires that two systems instantiate the same algorithm in the same architecture. For artificial and biological intelligence, this standard is neither attainable nor desirable. A transformer trained by gradient descent on text will never implement the same algorithm as a human brain, nor should it: the value of a plurality of intelligences lies precisely in their realizing cognition through different physical substrates and computational mechanisms. We therefore set it aside as an inappropriate target. At the other extreme, weak equivalence asks too little, requiring only matching input-output behavior while remaining agnostic about the processes that generate it. Between these extremes lies the grade we call strong equivalence: correspondence at the level of the seven process features. Two systems are process-equivalent to the extent that their behavior is constituted by the same process organization, even if that organization is implemented differently. This level of description is coarser than algorithms but finer than input-output mappings, and it is at this level that the properties of intelligence we care about reside. Process equivalence is therefore the standard adopted throughout the remainder of the paper. 3.3 A current illustration. This triparite version of the distinction matters in the current debate over what large language models represent. Binz and colleagues recently introduced Centaur 7, a model fine-tuned on over ten million human choices across hundreds of experiments, which predicts human behavior on held-out tasks better than bespoke models. Although its predictive power is impressive, the system is still only weakly equivalent. The model matches the outputs of human cognition without arriving at them as people do, and the match degrades under exactly the manipulations a process-level account would target, breaking down when small changes in wording shift meaning in ways human respondents track but the model does not 53. Inference from behavioral similarity/equivalence to shared underlying mechanisms is a well-known fallacy tracing back to Marr 37. Recent work explicitly recruits Marrâs levels framework to construe behavioral matching as a computational-level result, which underdetermines what is happening at both algorithmic- and implementation-levels 31. Centaur is the well-built limiting case of the trace-trained system, maximal in weak equivalence and silent on strong. The underlying concern is not new. Systems that match or exceed human performance without reproducing human process are a recurring occasion for this observation, from Deep Blueâs brute-force chess search 9 to earlier critiques of symbolic AI 15. 4 The machine case: a weak-equivalence engine 4.1 Trained on traces. Generative models are trained on corpora of textual and visual traces of human cognitive processes (e.g., finished papers, published code, transcripts, edited images). The iterative trials, abandoned approaches, dialogical pushback, and tacit-shaping behind these traces are largely absent from the corpora because they are rarely recorded. The structure of the training signal itself is the root of the process gap, yet this is not a shortage that more data can remedy. A larger corpus of traces is still a corpus of traces. Frontier models are, in this respect, unusually effective instruments of weak equivalence because they can reproduce traces of cognition without having been trained on the processes that generated them in humans. The question is, therefore, not whether scaling closes the process gap, but rather how much of the process itself can their agentic elaborations instantiate. 4.2 Cosmetic deliberation, and when it is not. Contemporary âreasoningâ and some agentic developments appear to add process: chain-of-thought externalizes intermediate steps 65; tree-of-thoughts branches over candidates 69; self-refinement and Reflexion add critique-and-revision loops 41; 56; coding agents run code, read failures, edit in response 67; and multi-agent debate subjects answers to challenge 16, all of which allocate inference-time compute to deliberation before answering 44; 22. Whether these scaffolds instantiate process, or only display and mimic it, is the question previous literature has examined. Consistently, reasoning traces routinely fail to track the computation driving an answer, larger models can be less faithful rather than more, contemporary reasoning systems reveal decision-shifting hints in fewer than one in five cases, and direct measurement finds models reward-hacking without verbalizing the hack 63; 32; 10; 29; 62. Debate scaffolds show a parallel pattern. Matched-compute comparisons find that multi-agent debate does not reliably beat single-agent self-consistency 57; 70. Debate over belief trajectories forms a martingale (i.e., on average each round of debate leaves the expected belief unchanged, so the exchanges add no information), implying that the intermediate exchanges contribute little beyond the eventual majority vote 11; 66. Reasoning-shaped output is, in these cases, additional output in the shape of reasoning rather than reasoning itself. This gap, however, is not intrinsic to the architecture. It follows from how models are trained and what tasks demand, and is therefore addressable by design. The same literature shows that the reasoning trace is not always cosmetic. When a task is difficult enough that serial computation is genuinely necessary (i.e., the answer cannot be reached without externalizing intermediate work), the reasoning trace becomes load-bearing in a computational sense. The answer is then computed through the externalized steps rather than alongside them, so a model cannot evade a monitor without abandoning the computation that produces the answer 17. Faithfulness is therefore a property of the task and of the training rather than of the architecture. A trace becomes faithful when the task forces the reasoning to do work, and it can be made more faithful by training that rewards verbalization 62. Across most of the currently deployed regimes, however, the task does not compel the reasoning to be load-bearing, and GenAI accordingly produces process-shaped output whose depicted process its architecture does not instantiate. 4.3 Feature-partial coverage. No existing scaffold covers all seven features at once. The features that models do not currently instantiate are the ones the human is left to supply (Fig. 2). Branching search and rejection sampling serve generative trial well, and self-refinement serves revision. However, the framing of the problem (what counts as a worthwhile question in the first place) still arrives from the prompt, so a person stays in-the-loop to perform the features an architecture lacks. Engagement with uncertainty is approximated by calibration and abstention, yet the system rarely distinguishes genuinely sitting in uncertainty from answering confidently, treating not-knowing as a threshold on when to answer rather than a mode of activity to remain in. The formative dimension is absent outright, since it plays out over a personâs development and no single system occupies that timescale. These gaps are not incidental. Better deliberation is pursued because it raises benchmark scores, not because theory says which features constitute process. Only such a theory can say which gaps matter and why. 4.4 Toward strong-equivalent architectures. If this gap is contingent on design, then closing it is a task of process engineering (i.e., building the generative activity into the architecture rather than optimizing its output). Our framework specifies what an architecture built for strong equivalence would have to do. We do not derive a full architecture from all seven features here. Instead, we develop two commitments concrete enough to build, each targeting a subset of the features rather than the whole set. The first is a dialogical-revision architecture (targeting dialogical accountability and value-laden framing), in which an agent revises because another agent has challenged how it framed the problem, not because a chain ran longer or a majority of agents agreed. Existing debate systems adjudicate on whether the agents end up agreeing, which is why they reduce to voting 16; 11. A framework-informed version would instead task the challenger with attacking the framing, rather than the answer, and reward substantive changes of mind over surface concessions. The second is a situated agentic loop (targeting feedback with the medium and framing), in which the environmentâs response can revise how the agent framed the task and not only the next action it takes. Coding agents are the nearest existing case, running code, reading failures, and editing in response, but their feedback updates the plan rather than the framing 68; 67. A frame-updating loop would separate a result that should change the plan from one that should change the problem itself, which requires holding the state of the problem apart from the candidate answer and a reward channel for reframing and sitting in uncertainty. Both commitments target the cycle directly, separating architectures that display it from those that run it. CoT ReAct ToT Self-Refine Reflexion Debate Reasoning models Coding agents Generative trial ~ ~ ⢠⢠⢠⢠~ ⢠Temporal extension ¡ ~ ¡ ¡ ~ ¡ ~ ~ Engagement w/ uncertainty ¡ ¡ ~ ~ ~ ~ ~ ¡ Feedback w/ medium ¡ ⢠¡ ~ ~ ¡ ¡ ⢠Value-laden framing ¡ ¡ ¡ ¡ ¡ ¡ ¡ ¡ Dialogical accountability ¡ ¡ ¡ ~ ~ ⢠¡ ¡ Formative dimension ¡ ¡ ¡ ¡ ¡ ¡ ¡ ¡ ⢠instantiated ~ displayed only ¡ structurally untouched Fig. 2: The machine process gap. Current agentic scaffolds (chain-of-thought, ReAct, tree-of-thoughts, self-refine, reflexion, multi-agent debate, reasoning models, coding agents) mapped against the seven process features, distinguishing features instantiated, features only displayed as process-shaped output, and features left structurally untouched. Coding agents instantiate generative trial and feedback with the medium through their runâfailâfix loop, yet leave value-laden framing, dialogue, and formation untouched, the features a human still supplies. 5 The human case: process on the human side 5.1 The same problem, on the human substrate. If the capacities we value are constituted in process, then a tool that performs the process on a personâs behalf does not simply save effort. It removes the activity through which the capacity would have formed, pushing human cognition toward the same weak equivalence that characterizes the machine. The cognitive-offloading literature establishes that what a person does with a tool changes what they can do without it 59; 51. Whether this is productive externalization or erosion, we argue, depends on whether the generative steps stay with the practitioner. On the human side, the seven features specify which parts of the activity a tool must leave intact. Where it absorbs the generative trial, sitting-in-uncertainty, dialogical challenge, or formative struggle, the humanâs own strong equivalence is at stake. 5.2 The evidence, and its boundary. The first wave of evidence is consistent enough to take seriously and careful enough to show where harm is and is not. A field experiment in secondary mathematics found that students with unrestricted GenAI access solved more problems while the tool was present but scored substantially worse once it was removed compared to unaided controls, the in-task gain having come from work the tool did for them 4. The same pattern recurs across a range of process-level measures, including poorer performance on delayed knowledge-retention tests following chatbot-assisted study 3, weaker neural connectivity and poorer recall of oneâs own just-written text after assisted essay-writing 30, reduced metacognitive engagement (âmetacognitive lazinessâ) 19, and lower mental effort at the cost of depth in scientific inquiry 60. The cost extends to expertise. For instance, adopting AI-assisted colonoscopy procedures reduced cliniciansâ unaided detection rates, suggesting deskilling effects among practitioners who were previously proficient 8. The cost can also go unfelt. In a randomized trial, experienced open-source developers completed tasks more slowly with AI assistance than without it, while believing they were faster, suggesting even expert practitioners may not reliably perceive the toolâs effects 5. The convergence is not yet a causal demonstration of long-term formative effects, but the short-term pattern is what our framework predicts. Assistance that performs the process a person would otherwise have carried out forecloses the capacity that process builds. However, the evidence does not support a blanket ban on such tools. The outcome depends on the role the tool plays, and the boundary condition is whether the tool performs the generative step or leaves it with the person. The same mathematics experiment found that a version constrained to give hints and withhold answers preserved learning where the unconstrained version harmed it, locating the harm not in AI assistance itself but in whether the generative step is withheld from the learner 4. A randomized study of adults learning a new programming library found the same dividing line. Interactions that kept the user cognitively engaged preserved learning, whereas full delegation produced productivity without competence 55. The essay-writing study makes the point from the other direction. Participants who wrote unaided first and only then turned to the tool showed greater neural connectivity than those who used it throughout, the assistance arriving after the generative work rather than in place of it 30. The strongest counter-evidence shows students taught by a well-designed AI tutor learned more, and in less time, than peers in an active-learning classroom 28. Notably, the tutor was built to keep students doing the generative reasoning rather than to deliver solutions, and gains were measured immediately rather than on delayed, unaided transfer tests, which may expose erosion elsewhere. Whether such scaffolded gains persist once the tutor is removed is the question a process-level account makes central. Indeed, process-preserving and process-substituting assistance are different constructs, and the seven features distinguish them. 5.3 The formative stakes. What is at stake in the erosive case is not only skill but formation. The capacities the process account foregrounds (judgment, recognition of ill-posed problems, taste that discriminates a promising direction, practical wisdom the tradition calls phronesis) are formed slowly and below articulation, through exactly the difficult, uncertain, dialogical activity that capable assistants make most tempting to offload 1; 15. A tool that reliably supplies the answer removes the occasions on which these capacities would have been exercised and formed, and does so most efficiently for the learner who most needs the practice. The risk is developmental and cumulative, falling hardest on the next generation of practitioners. 5.4 Process-preserving design. If process is what cognition is constituted in, the design principle follows directly. It is not a matter of less assistance or more friction but of which parts of the activity the tool keeps with the person. An assistant that supplies the answer compresses the generative trial, while one that supplies the next question or counterexample extends it. An assistant that resolves an ambiguity dissolves the uncertainty the user should have sat with, while one that raises it and declines to settle returns the engagement to the user. An assistant that capitulates flatters, while one that holds a well-grounded position supplies the dialogical challenge the process requires (Table 2). These are the same features that specify strong-equivalent machine architectures, now read as constraints on the humanâAI loop rather than on the agent alone. The two design programs converge under a single standard of process. What the field lacks is not the principles but the comparative evidence (i.e., studies testing process-preserving against process-substituting assistance on the downstream, unaided capacities that matter). Table 2: Process-preserving versus process-substituting assistance. The same seven features that diagnose machine cognition specify, on the human side, which parts of the activity a tool must leave with the person, contrasting a process-preserving and a process-substituting assistant feature by feature. Process feature Process-preserving assistant Process-substituting assistant Generative trial Supplies the next question or counterexample Supplies the answer; the trial is compressed Temporal extension Spreads the work across returns Collapses the work to a single shot Engagement w/ uncertainty Surfaces the ambiguity, declines to settle it Resolves the ambiguity confidently Feedback w/ medium Keeps the user in contact with the material Mediates the material away Value-laden framing Leaves the framing of the problem to the user Fixes the framing in advance Dialogical accountability Holds a well-grounded position Capitulates and flatters Formative dimension Retains the occasions on which judgment forms Removes them; the capacity goes unbuilt 6 A shared measurement via process audits The framework matters only if the seven features can be measured, and pitching strong equivalence at the resolution of the features rather than the algorithm is what makes shared measurement possible. A process audit is a task-and-rubric protocol that scores a solverâs behavioral trace against the seven features, with substrate-specific behavioral anchors. Because the features are defined at a coarser resolution than the algorithm, the same audit can be administered to human and machine solvers on matched tasks. Conventional evaluation scores only the output. On reasoning items it rewards the correct answer, on creative tasks it rates the output. A process audit instead scores the activity (whether the solver generated and revised alternatives, flagged an ill-posedness rather than answering through it, updated on feedback, and marked the framing of the problem as a choice), so that two solvers reaching the same answer can receive different scores, and a solver reaching a worse answer through richer process can score higher on the dimensions the framework says matter. Work in comparative cognition calls for exactly this, probing machine cognition with the process-sensitive methods developed for animal and human cognition rather than output benchmarks alone 27; 64; 48. Several probes make specific features measurable (Table 3). An ill-posed reasoning probe presents an under-determined problem with a confidently signaled expected answer and scores whether the solver flags the ill-posedness, develops alternative framings, and refrains from premature closure. A mid-trace perturbation, of the kind the faithfulness literature already employs, injects new information partway through a solution and scores whether the trace genuinely revises or merely continues 32. A dialogical-accountability probe scores whether a post-challenge trajectory engages the substance of a challenge rather than its chain length. Each yields a behavioral signature scorable on a human transcript and a machine trace alike, letting the audit compare process across substrates rather than within one. A shared audit turns the frameworkâs central commitments into empirical claims. On the machine side, an architecture that implements a feature should score measurably higher on that featureâs audit than the scaffold it replaces, with the gain intact after controlling for output quality. An intervention that raises benchmark scores while leaving audit scores unchanged is improving output, not process. On the human side, the audit supplies the missing dependent measure for the comparative studies called for above, namely whether process-preserving assistance, against process-substituting assistance, protects downstream capacities (unaided transfer, calibrated uncertainty, recognition of ill-posed problems, quality of independent revision) that the short-term evidence places at risk. Whether feature-targeted architectures close the machine-side gap, and which preserved features protect which human capacities, are open questions the framework makes precise and the audit makes measurable. Table 3: Process audits, with the probes, the features they target, and the behavioral signatures scored. (Administrable to human and machine solvers on matched tasks.) Probe Features targeted Behavioral signature scored Ill-posed problem Engagement with uncertainty; Value-laden framing Flags ill-posedness; develops alternative framings; refuses premature closure Mid-trace perturbation Generative trial and revision; Feedback with the medium Genuinely revises vs continues unchanged after injected information Dialogical challenge Social and dialogical accountability; Value-laden framing Engages the substance of a challenge vs concedes or pads chain length Sustained / return task Temporal extension; Formative dimension Maintains and updates state across returns; carries learning forward (longitudinal on the human side) 7 Conclusion Intelligence is constituted in process, and the distinction drawn above (weak equivalence at the output, strong equivalence at the resolution of process) separates a system that matches cognition from one that instantiates it. Current GenAI, trained on the traces of human cognition, is built for the first and largely silent on the second. Faithfulness evidence shows how far reasoning-shaped output can diverge from reasoning-constituting activity wherever tasks do not force the reasoning to do work. Distinguishing strong from hard equivalence lets the same diagnosis apply without requiring that machines reproduce human algorithms and keeps a plurality of intelligences in view. The cost of ignoring the features is paid twice, by machine systems that produce the shape of cognition without its substance and by human users whose own process is left uncultivated. The constructive consequence is that the same criterion does work on both sides. On the machine side it specifies architectures (answerable to challenge, situated in a medium that can overturn a framing) that instantiate more of the process rather than more of its appearance. On the human side, it specifies tools that preserve the generative, uncertain, dialogical, and formative activity through which judgment is built. Process audits make the standard measurable on both, turning its commitments into answerable questions, namely whether architectures designed against the features score higher without cosmetic gain and which preserved features protect which human capacities under sustained assistance. A useful companion to this framing has been the rise of prompt and context engineering, which improves what systems return by optimizing how they are conditioned. Process engineering, in the sense developed above, points past it toward the generative activity itself rather than the conditioning of its output. The distinction bears on the largest question in view. If the capacities we call general are constituted in process rather than read off a distribution of outputs, then optimizing outputs alone may approximate general intelligence without constituting it, a conjecture the process audits are designed to test rather than a prediction we are in a position to make. What intelligence is, and what it is for, are old questions. What is new is that systems producing its outputs are now built and deployed at scale, which makes the difference between an intelligenceâs outputs and its process consequential in a way it has never been before. Acknowledgements This work was supported by a Macquarie University 2025 Innovation in Education Grant (Spark) and by an Advanced Strategic Capabilities Accelerator (ASCA) grant (AN-12973) from the Australian Department of Defence, in collaboration with the Defence Science and Technology Group (DSTG). Use of generative AI The authors used generative AI tools to assist with editing, concision, and revision. The intellectual content is the authorsâ own, and the authors reviewed and take full responsibility for the content of the publication. Author contributions M.J.R. and A.A. conceived and led the work, and M.J.R. wrote the first draft of the manuscript. A.A., C.C., M.P.D.M., and P.N. contributed substantially to developing the ideas and to drafting and revising the manuscript. M.D., R.W.K., and D.K. contributed to refining the framework and revising the manuscript. All authors reviewed and approved the final manuscript. Disclosure of interests The authors have no competing interests to declare that are relevant to the content of this article. Data availability This is a theoretical study; no datasets were generated or analysed, and no custom code was produced. References Aristotle (2009) The nicomachean ethics. Oxford University Press. Note: L. Brown, Ed., D. Ross, Trans. Cited by: §5. M. M. Bakhtin (1981) The dialogic imagination: four essays. University of Texas Press. Cited by: §2. A. Barcaui (2025) ChatGPT as a cognitive crutch: evidence from a randomized controlled trial on knowledge retention. Social Sciences & Humanities Open 12 (), p. 102287. Note: 10.1016/j.ssaho.2025.102287 Cited by: §5. H. Bastani, O. Bastani, A. Sungu, H. Ge, Ă. KabakcÄą, and R. Mariman (2025) Generative AI without guardrails can harm learning: evidence from high school mathematics. Proceedings of the National Academy of Sciences 122 (26), p. e2422633122. Note: 10.1073/pnas.2422633122 Cited by: §1, §5, §5. J. Becker, N. Rush, E. Barnes, and D. Rein (2025) Measuring the impact of early-2025 AI on experienced open-source developer productivity. Technical report METR. Note: arXiv:2507.09089 Cited by: §5. E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell (2021) On the dangers of stochastic parrots: can language models be too big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT â21), p. 610â623. External Links: Document Cited by: §1. M. Binz et al. (2025) A foundation model to predict and capture human cognition. Nature 644, p. 1002â1009. Note: 10.1038/s41586-025-09215-4 Cited by: §3. K. BudzyĹ et al. (2025) Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. The Lancet Gastroenterology & Hepatology 10 (10), p. 896â903. Note: 10.1016/S2468-1253(25)00133-5 Cited by: §5. M. Campbell, A. J. Hoane, and F. Hsu (2002) Deep blue. Artificial Intelligence 134 (1â2), p. 57â83. Cited by: §3. Y. Chen et al. (2025) Reasoning models donât always say what they think. arXiv:2505.05410. Note: 10.48550/arXiv.2505.05410 Cited by: §1, §4. H. K. Choi, J. Zhu, and S. Li (2025) Debate or vote: which yields better decisions in multi-agent large language models?. In Advances in Neural Information Processing Systems, Vol. 38, p. 101732â101764. Note: arXiv:2508.17536 Cited by: §4, §4. I. D. Couzin (2009) Collective cognition in animal groups. Trends in Cognitive Sciences 13 (1), p. 36â43. Note: 10.1016/j.tics.2008.10.002 Cited by: §2. M. R. W. Dawson (2013) Mind, body, world: foundations of cognitive science. Athabasca University Press. Cited by: §3. J. Dewey (1938) Logic: the theory of inquiry. Holt. Cited by: §2. H. L. Dreyfus and S. E. Dreyfus (1986) Mind over machine: the power of human intuition and expertise in the era of the computer. The Free Press. Cited by: §3, §5. Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch (2024) Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning, p. 11733â11763. Cited by: §1, §4, §4. S. Emmons et al. (2025) When chain-of-thought is necessary, language models struggle to evade monitors. arXiv:2507.05246. Note: 10.48550/arXiv.2507.05246 Cited by: §4. K. A. Ericsson, R. T. Krampe, and C. Tesch-RĂśmer (1993) The role of deliberate practice in the acquisition of expert performance. Psychological Review 100 (3), p. 363â406. Note: 10.1037/0033-295X.100.3.363 Cited by: §2, §2. Y. Fan et al. (2025) Beware of metacognitive laziness: effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology 56 (2), p. 489â530. Note: 10.1111/bjet.13544 Cited by: §5. J. A. Fodor and Z. W. Pylyshyn (1988) Connectionism and cognitive architecture: a critical analysis. Cognition 28 (1â2), p. 3â71. Note: 10.1016/0010-0277(88)90031-5 Cited by: §3. D. M. Gordon (2010) Ant encounters: interaction networks and colony behavior. Princeton University Press. Cited by: §2. D. Guo et al. (2025) DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, p. 633â638. Note: 10.1038/s41586-025-09422-z Cited by: §1, §4. J. Hadamard (1945) An essay on the psychology of invention in the mathematical field. Princeton University Press. Cited by: §1, §2, §2. A. M. Haith (2026) Policy-gradient reinforcement learning as a general theory of practice-based motor skill learning. eLife. Note: Reviewed preprint External Links: Document Cited by: §2. J. Haugeland (1998) Having thought: essays in the metaphysics of mind. Harvard University Press, Cambridge, MA. Cited by: §2. T. Ingold (2013) Making: anthropology, archaeology, art and architecture. Routledge. Cited by: §2. A. A. Ivanova (2025) How to evaluate the cognitive abilities of LLMs. Nature Human Behaviour 9, p. 230â233. Note: 10.1038/s41562-024-02096-z Cited by: §6. G. Kestin, K. Miller, A. Klales, T. Milbourne, and G. Ponti (2025) AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports 15, p. 17458. Note: 10.1038/s41598-025-97652-6 Cited by: §5. T. Korbak et al. (2025) Chain of thought monitorability: a new and fragile opportunity for AI safety. arXiv:2507.11473. Note: 10.48550/arXiv.2507.11473 Cited by: §1, §4. N. Kosmyna et al. (2025) Your brain on ChatGPT: accumulation of cognitive debt when using an AI assistant for essay writing task. arXiv:2506.08872. Note: 10.48550/arXiv.2506.08872 Cited by: §5, §5. A. Ku et al. (2025) Levels of analysis for large language models. Philosophical Transactions of the Royal Society A 384 (2320), p. 20250012. Note: 10.1098/rsta.2025.0012 Cited by: §3. T. Lanham et al. (2023) Measuring faithfulness in chain-of-thought reasoning. arXiv:2307.13702. Note: 10.48550/arXiv.2307.13702 Cited by: §1, §4, §6. J. Lave and E. Wenger (1991) Situated learning: legitimate peripheral participation. Cambridge University Press. Cited by: §2, §2. Y. LeCun (2022) A path towards autonomous machine intelligence. Technical report OpenReview. Note: Position paper, version 0.9.2 Cited by: §2. H. Lee et al. (2025) The impact of generative AI on critical thinking: self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, p. 1â22. Note: 10.1145/3706598.3713778 Cited by: §1. M. Levin (2022) Technological approach to mind everywhere: an experimentally-grounded framework for understanding diverse bodies and minds. Frontiers in Systems Neuroscience 16, p. 768201. Note: 10.3389/fnsys.2022.768201 Cited by: §2. Z. Lin (2025) Six fallacies in substituting large language models for human participants. Advances in Methods and Practices in Psychological Science. Note: 10.1177/25152459251357566 Cited by: §3. H. E. Longino (1990) Science as social knowledge: values and objectivity in scientific inquiry. Princeton University Press. Note: 10.2307/j.ctvx5wbfz Cited by: §2. P. Lyon (2015) The cognitive cell: bacterial behavior reconsidered. Frontiers in Microbiology 6, p. 264. Note: 10.3389/fmicb.2015.00264 Cited by: §2. A. MacIntyre (1981) After virtue: a study in moral theory. University of Notre Dame Press. Cited by: §2. A. Madaan et al. (2023) SELF-REFINE: iterative refinement with self-feedback. In Proceedings of the 37th International Conference on Neural Information Processing Systems, p. 46534â46594. Cited by: §1, §4. D. Marr (1982) Vision: a computational investigation into the human representation and processing of visual information. Holt. Cited by: §3. T. Nakagaki, H. Yamada, and Ă. TĂłth (2000) Maze-solving by an amoeboid organism. Nature 407, p. 470. Note: 10.1038/35035159 Cited by: §2. OpenAI (2024) OpenAI o1 system card. Technical report OpenAI. Note: arXiv:2412.16720 Cited by: §1, §4. M. Polanyi (1958) Personal knowledge: towards a post-critical philosophy. University of Chicago Press. Cited by: §1, §2. M. Polanyi (1966) The tacit dimension. Doubleday. Cited by: §2. Z. W. Pylyshyn (1984) Computation and cognition: toward a foundation for cognitive science. MIT Press. Note: 10.7551/mitpress/2004.001.0001 Cited by: §1, §1, Fig. 1, §3. S. Rane et al. (2025) Position: principles of animal cognition to improve LLM evaluations. In Proceedings of the 42nd International Conference on Machine Learning, p. 82051â82061. Cited by: §6. C. R. Reid, T. Latty, A. Dussutour, and M. Beekman (2012) Slime mold uses an externalized spatial "memory" to navigate in complex environments. Proceedings of the National Academy of Sciences 109 (43), p. 17490â17494. Note: 10.1073/pnas.1215037109 Cited by: §2. M. J. Richardson, M. P. Diaz Monfort, A. Alhasan, C. Crone, S. Tyagi, M. Varlet, M. Dras, and R. W. Kallen (2026) Process, not trace: a theoretical framework for agentic AI and process-preserving humanâAI cognition. In Advances in Practical Applications of Agentic AI and Multi-Agent Systems: The PAAMS Collection, Lecture Notes in Computer Science. Note: In press Cited by: §1, footnote. E. F. Risko and S. J. Gilbert (2016) Cognitive offloading. Trends in Cognitive Sciences 20 (9), p. 676â688. Note: 10.1016/j.tics.2016.07.002 Cited by: §5. D. A. SchĂśn (1983) The reflective practitioner: how professionals think in action. Basic Books. Cited by: §1, §2, §2. S. SchrĂśder, T. Morgenroth, U. Kuhl, V. Vaquet, and B. PaaĂen (2025) Large language models do not simulate human psychology. arXiv:2508.06950. Note: 10.48550/arXiv.2508.06950 Cited by: §3. T. D. Seeley (2010) Honeybee democracy. Princeton University Press. Cited by: §2. J. H. Shen and A. Tamkin (2026) How AI impacts skill formation. arXiv:2601.20245. Note: 10.48550/arXiv.2601.20245 Cited by: §5. N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao (2023) Reflexion: language agents with verbal reinforcement learning. In Proceedings of the 37th International Conference on Neural Information Processing Systems, p. 8634â8652. Cited by: §4. A. P. Smit, N. Grinsztajn, P. Duckworth, T. D. Barrett, and A. Pretorius (2024) Should we be going MAD? a look at multi-agent debate strategies for LLMs. In Proceedings of the 41st International Conference on Machine Learning, p. 45883â45905. Cited by: §4. B. C. Smith (2019) The promise of artificial intelligence: reckoning and judgment. MIT Press, Cambridge, MA. Cited by: §2. B. Sparrow, J. Liu, and D. M. Wegner (2011) Google effects on memory: cognitive consequences of having information at our fingertips. Science 333 (6043), p. 776â778. Note: 10.1126/science.1207745 Cited by: §5. M. Stadler, M. Bannert, and M. Sailer (2024) Cognitive ease at a cost: LLMs reduce mental effort but compromise depth in student scientific inquiry. Computers in Human Behavior 160, p. 108386. Note: 10.1016/j.chb.2024.108386 Cited by: §5. A. Tero et al. (2010) Rules for biologically inspired adaptive network design. Science 327 (5964), p. 439â442. Note: 10.1126/science.1177894 Cited by: §2. M. Turpin, A. Arditi, M. Li, J. Benton, and J. Michael (2025) Teaching models to verbalize reward hacking in chain-of-thought reasoning. arXiv:2506.22777. Note: 10.48550/arXiv.2506.22777 Cited by: §4, §4. M. Turpin, J. Michael, E. Perez, and S. R. Bowman (2023) Language models donât always say what they think: unfaithful explanations in chain-of-thought prompting. In Proceedings of the 37th International Conference on Neural Information Processing Systems, p. 74952â74965. Cited by: §1, §4. K. Voudouris, L. Cheke, and E. Schulz (2025) Bringing comparative cognition approaches to AI systems. Nature Reviews Psychology 4, p. 363â364. Note: 10.1038/s44159-025-00456-8 Cited by: §6. J. Wei et al. (2022) Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, p. 24824â24837. Cited by: §1, §4. H. Wu, Z. Li, and L. Li (2025) Can LLM agents really debate? a controlled study of multi-agent debate in logical reasoning. arXiv:2511.07784. Note: 10.48550/arXiv.2511.07784 Cited by: §4. J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press (2024) SWE-agent: agent-computer interfaces enable automated software engineering. In Advances in Neural Information Processing Systems 37 (NeurIPS), Cited by: §4, §4. S. Yao et al. (2023a) ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), Note: 10.48550/arXiv.2210.03629 Cited by: §1, §4. S. Yao et al. (2023b) Tree of thoughts: deliberate problem solving with large language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems, p. 11809â11822. Cited by: §1, §4. H. Zhang et al. (2025) Position: stop overvaluing multi-agent debate â we must rethink evaluation and embrace model heterogeneity. arXiv:2502.08788. Note: 10.48550/arXiv.2502.08788 Cited by: §4. Supplementary Information This supplement expands the seven features of process-constituted intelligence introduced in the main text (Table 1 gives compressed biological, human, and machine instantiations). For each feature, Supplementary Table 1 states a working definition, a human and an AI/machine instantiation, and a diagnostic signatureâthe observable evidence an auditor would look for to judge the feature present rather than absent. The diagnostic column is the operational bridge between the seven features and the process audits of Section 6: it says, feature by feature, what evidence would distinguish process from a trace that merely resembles it. Supplementary Table 1: Seven features of process-constituted intelligence: a working definition, a human and an AI/machine instantiation, and a diagnostic signature (what an observer looks for to judge each feature present versus imitated; maps onto the Section 6 process audits). Feature Definition Human instantiation AI / machine instantiation Diagnostic signature (cf. S6) 1. Generative trial and revision Cognition proceeds by producing many candidate attempts and refining them through their failures; the discarded attempts are the substance of the work, not waste preliminary to it. An inventorâs failed prototypes; a mathematicianâs abandoned proof strategies; a painterâs discarded studies; drafting and redrafting an email, or trying several wordings until one fits. Multi-sample rollouts, tree-of-thought branching, rejection sampling with self-critique; a coding agentâs run-fail-fix loop against a test suite. Are failed candidates actually generated, retained, and shown to shape the final output or is a single forward pass presented as if it were a search? Look for a traceable revision history, not just a polished result. 2. Temporal extension Cognition unfolds over time and across repeated returns to the same material, accumulating state rather than resolving in one bounded step. A scientistâs months on a single problem; the years of deliberate practice that build a skill; mulling a hard decision over several days; or learning to cook or drive over months. Persistent agentic trajectories that carry state across sessions; memory that accrues across interactions rather than resetting each prompt. Does earlier work measurably constrain later work, with state maintained and revisited across turns or is each response stateless and bounded by a single context window? 3. Engagement with uncertainty A solver sits in not-knowing, recognizes when a problem is ill-posed, and refuses premature closure rather than resolving to confidence the situation does not warrant. Recognizing that a question is ill-posed; withholding judgment until the evidence warrants it; saying âIâm not sure yetâ and seeking more information before deciding. Calibrated probabilities; explicit ambiguity flagging; abstention or clarification-seeking in place of confident confabulation. On under-specified or unanswerable inputs, does the system flag, abstain, or ask or produce a fluent, confident answer regardless? Probe with ill-posed prompts and measure calibration. 4. Feedback with the medium The material (e.g., proof, canvas, instrument, dataset, codebase) talks back and redirects the activity, so the work is a conversation with the medium rather than execution of a pre-formed plan. A painter responding to what the canvas does; a mathematician led by what the proof will and will not permit; adjusting a recipe as you taste it; or rearranging a room by seeing how it looks. Tool-using agents that update their framing on environmental return (compiler errors, test failures, retrieved evidence), not merely their next token. Does environmental return change the plan and framing, or only the surface output? Distinguish genuine reframing from re-running the same plan against feedback. (Coding agents come closest; S4.) 5. Value-laden framing What counts as a promising move or a worthwhile problem is constituted by the practitionerâs developing judgment and is itself cognitive work, not a parameter fixed in advance. Expert judgment of which problems are worth posing and which moves are promising; sensing which task on a busy day actually matters, or which point in a disagreement is the real one. Objectives or reward-shaping that target problem-reframing rather than solution-pursuit; in current systems the framing is largely supplied from outside. Does the system generate or revise its own framing of what matters, or optimize a framing handed to it? Look for reframing of the goal, not just efficient pursuit of a fixed one. 6. Social and dialogical accountability Cognition is constituted in part by being answerable to others: reasoning with and against interlocutors and having oneâs framings challenged. Practices built around being answerable to others: chavruta (paired study in which partners argue a text into clarity), Socratic dialectic (a claim tested through question and counter-question), and peer review (claims certified only after expert scrutiny); explaining your reasoning to a friend who pushes back; or defending a plan to colleagues who challenge it. Multi-agent adversarial revision; inter-agent challenge in which a critic can alter the framing rather than merely vote on the output. Can adversarial critique change the framing and the outcome or is the âdialogueâ cosmetic, with agents that concur or rubber-stamp? Test whether challenge measurably alters results. 7. Formative dimension Much of the process runs below explicit articulation, is built up through embodied practice within a community, and simultaneously shapes who the practitioner is becoming. Becoming the kind of scientist or artist who can do this work; tacit skill acquired through sustained practice; growing into a parentâs judgment, or a driverâs feel for the road. Beyond single-system scope; appears, if at all, through humanâAI co-formation (S5); the system shapes, and is shaped within, a practice rather than internalizing one itself. Is there development of tacit, practice-grounded capacity over time, or fixed weights invoked per task? For current systems the honest answer is largely absent; assess at the humanâAI system level (S5).