Paper deep dive
Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World
Jiang Jiang, Yifu Sun, Qi Shen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/18/2026, 5:47:16 AM
Summary
The paper proposes a 'legitimacy theory' for constitutive priors in machine intelligence, arguing that physical AI is stalled by a 'cold-start deadlock' where data requires intelligence and intelligence requires data. The authors identify the 'artificial physical world' (buildings, infrastructure) as an exception where prior frameworks (design archives) exist before instances. They derive a four-world ontology, a legitimacy criterion for prior extraction based on intentional constitution and readable archives, and a four-layer framework (syntax, concept, knowledge, instance). The paper stakes five falsifiable predictions and positions Large Language Models as readers of these archives rather than the archives themselves.
Entities (10)
Relation Signals (8)
Artificial Physical World → hasarchive → Readable Archive
confidence 95% · designed artifacts ship with readable archives — drawings, manuals, operating procedures — that precede their instances
Artificial Physical World → hasproperty → Intentional Constitution
confidence 95% · buildings, industrial facilities, and infrastructure form the only class that is both intentionally constituted and documented
Constitutive Priors → islegitimatein → Artificial Physical World
confidence 94% · prior extraction is legitimate if and only if the object domain is intentionally constituted and has left a readable archive
Constitutive Priors → solves → cold-start deadlock
confidence 93% · a prior framework gives intelligence that runs from day one; running itself produces data; data flows back to refine the model; the deadlock becomes a flywheel
Constitutive Priors → requiresstructure → Four-Layer Framework
confidence 92% · any such framework has at least four layers -- syntax, concept, knowledge, instance
Large Language Models → rolein → Constitutive Priors Framework
confidence 90% · Large language models find an honored place here -- as readers of the archive, not as the archive
Physics-Informed Neural Networks → iscontrolgroupfor → Constitutive Priors
confidence 85% · with physics-informed neural networks (PINNs) as the control group
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Machine intelligence has conquered the symbolic world but stalled at the physical one. The stall is structural: physical AI faces a cold-start deadlock -- no intelligence without data, no data without deployed intelligence. Our thesis: the deadlock is real but unevenly distributed, and the exception has a name: the artificial physical world. Buildings, industrial facilities, and infrastructure are intentionally constituted and documented: designed artifacts ship with readable archives that precede and constitute their instances; here, norms are promulgated before instances, not averaged from them. Four contributions. (i) From a four-world ontology we derive a legitimacy criterion for constitutive prior frameworks: prior extraction is legitimate if and only if the object domain is intentionally constituted and has left a readable archive; the criterion is testable through direction of fit -- deviation from a constitutive norm is a violation in the world, not a revision of the model. (ii) We establish a layering lower bound: any such framework has at least four layers -- syntax, concept, knowledge, instance -- because four construction goals pair into mutually incompatible carriers. (iii) We register deployment claims across five industrial domains and a 32-class failure-mode vocabulary. (iv) We stake the framework on five falsifiable predictions, the central one checkable on the public engineering record: if it fails, the framework fails. Semi-formal arguments back these claims (Appendix A): a Gold-type boundary on rule coverage in archiveless worlds, a decidability result for failure reduction over closed concept layers, and a boundary theorem for certificate-anchored calculi. Large language models find an honored place here -- as readers of the archive, not as the archive. First of three companion works; the companions take up the questions deliberately left open.
Tags
Links
- Source: https://arxiv.org/abs/2608.15147v1
- Canonical: https://arxiv.org/abs/2608.15147v1
Trouble viewing inline? Open PDF directly →
Full Text
203,168 characters extracted from source content.
Expand or collapse full text
Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World Jiang Jiang Thanks: Corresponding author: jiangjiang@persagy.com Yifu Sun Qi Shen Affiliation: [2pt] Persagy Science and Technology Co., Beijing, China Affiliation: jiangjiang@persagy.com, sunyifu@persagy.com, shenqi@persagy.com August 2026 Abstract Machine intelligence has conquered the symbolic world but stalled at the physical one — and the stall is structural: physical AI faces a cold-start deadlock, in which no intelligence can be trained without data, and no data is produced without deployed intelligence. Our thesis: the deadlock is real but unevenly distributed — and the exception has a name: the artificial physical world. Among the physical worlds considered here, buildings, industrial facilities, and infrastructure form the only class that is both intentionally constituted and documented: designed artifacts ship with readable archives — drawings, manuals, operating procedures — that precede their instances and constitute them. Here, norms are promulgated before instances, not averaged from them. Four contributions. (i) Building on a four-world ontology (the symbolic world, the artificial physical world, the basic physical world, the phenomenal world), we derive a legitimacy criterion for constitutive prior frameworks: prior extraction is legitimate if and only if the object domain is intentionally constituted and has left a readable archive; the criterion is testable through direction of fit — deviation from a constitutive norm is a violation in the world, not a revision of the model; the four-world partition itself is then used, in Section 8, to take the measure of an independently proposed theory. (i) We establish a layering lower bound: any such framework must have at least four layers — syntax, concept, knowledge, instance — because four construction goals pair into mutually incompatible carriers. (i) We register the framework’s deployment claims and cross-domain commitments across five industrial domains, together with a 32-class enumerative vocabulary of failure modes. (iv) We stake the framework on five falsifiable predictions, together with publicly specified tests for evaluating them — the central prediction (cross-industry invariance of the concept layer) is checkable on the public engineering record: if the prediction fails, the framework fails. Semi-formal arguments back these claims (Appendix A): a Gold-type boundary on rule coverage in archiveless worlds, a decidability result for failure reduction over closed concept layers, and an explicit boundary theorem for the applicability domain of certificate-anchored calculi. Large language models find an honored place in this architecture — as readers of the archive, not as the archive. The paper is the first of three companion works; the two companions take up the questions deliberately left open here: the form of the intelligent body, and the carrier of the framework’s knowledge and instance layers. Keywords: constitutive priors ⋅· prior knowledge ⋅· legitimacy ⋅· artificial physical world ⋅· layered ontology ⋅· direction of fit ⋅· falsifiable predictions 1. Introduction: An Overlooked Asymmetry The most expensive lesson in the history of twentieth-century artificial intelligence can be compressed into one sentence: build the framework first, learn later — that road is closed. Rule-based NLP erected the framework from grammar books and then asked machines to learn language by it --- it died. Expert systems erected the framework in knowledge-engineering interviews and then asked rule bases to reason --- they died. Cyc [1], at the cost of roughly a person-century, tried to hand-write human common sense into a framework in its entirety --- it died too. Statistical learning swept the field with ‘abandon the prior framework’11 1 Terminology. The “prior” of this paper differs from Kant’s a priori. Kant’s categories are the innate forms that the subject legislates to experience — they need no archive, being built into the knower; our prior framework is world-side: promulgated by intentional agents, external to any learner in documentary form, and a posteriori extractable. The rationalism–empiricism dispute presupposed a dilemma — prior knowledge is either innate or nonexistent; this paper points to a third road: the prior can be externalized, documented, and extractable. The full demarcation is in §3.4. as its ticket of admission, and the very word framework became nearly a slur inside machine learning, roughly synonymous with “top-down naivety.” Yet over the same stretch of history, the most expensive artifacts of human civilization took precisely the opposite route. Every building has drawings before construction; every power station has control logic before grid connection; every aircraft has design specifications before airworthiness certification. The engineering world does not merely tolerate “framework first, operation later” — it is that, and it works so well that we entrust our lives to it: the floor slab above your head, the grid beneath your feet, the flight outside your window are all instances of prior frameworks. The same act that is a capital crime in language is daily routine in engineering. The asymmetry is enormous, yet it has rarely been studied as a problem in need of explanation. Our basic position: the line between the dead and the living does not run between methods, but between worlds. An old question has become urgent. The breakthrough of large language models has given the entire industry a collective impatience: AI has proved its power in the symbolic world, and everyone is eager to wire it into the real physical one — only to find that it does not connect. So the search for detours began: massive simulation at the phenomenal layer, attempts to reverse-engineer the causal structure of the physical world from synthetic data and oceanic interaction — investments in the tens or hundreds of billions of dollars; “the GPT moment of embodied AI” is announced every few weeks — while the robots under the spotlight dance, spar, and pick boxes in factories. Not for lack of effort; but something always seems to stand in between. And what stands in between is, in fact, the distribution of data. Every breakthrough of machine learning in the past decade happened in the symbolic world — text, code, images — where trillion-token corpora exist: free, standardized, written to be read. The physical world has no such corpus. Take buildings and industrial facilities: every building is a unique instance — no two share a configuration; operational telemetry is the owner’s private property, scattered across disconnected building-automation systems; and the most informative data of all — failure data — is both rare and expensive, since no owner will let you run a building into failure to collect a training set. Physical AI is thus caught in a cold-start deadlock: no intelligence without data, no data without deployed intelligence. Neither mainstream route unties the knot head-on: the LLM route stays in the symbolic world, speaking of the physical only indirectly through language; the world-model and simulation route tries to simulate everything at the phenomenal layer, trading generality for data — but what it simulates is the phenomenal world, which has no archive, whereas the place where robots will ultimately work, and where asset value resides, is the artificial physical world, which has one. This paper proposes a third road: if some physical world permits erecting the framework first, the deadlock can be untied at its source — a prior framework gives intelligence that runs from day one; running itself produces data; data flows back to refine the model; the deadlock becomes a flywheel. One question remains: which world permits it? Research question. Given a domain W, under what conditions may an artificial intelligence system legitimately22 2 “Legitimacy” is a technical usage here, unrelated to the normative notions of philosophy: it names the conditions under which a machine-learning system is justified in extracting a prior framework from a domain’s archives. We make no claim, and need none, about epistemic, moral, or political legitimacy in the philosophical literature (§4.2). acquire a prior framework33 3 Throughout, “prior” is used in a non-Bayesian sense. A Bayesian prior is a probability distribution over parameters or hypotheses — elicited from belief, born to be revised by data; a constitutive prior is a promulgated normative framework — extracted from archives, born to constrain data. The full demarcation is in Section 8, R.2’. before accumulating large-scale observational data? The question decomposes into three: (P1) Under what conditions does an extracted framework constrain instances as promulgated norm rather than statistical average — so that failure becomes violation, loud and locatable? (P2) What shape must such a framework take, in order to remain one coherent system while the world it covers keeps growing? (P3) How can these two answers be tested publicly — falsifiably, on the record? All three conditions issue from a single source: prediction is by construction description; correctness is promulgated. That, then, is the research question of this paper: when, where, and for which worlds is a prior framework legitimate for machine learning? Note that we do not ask “prior or posterior, which is better” — that is a false opposition, and history long ago proved the answer domain-dependent. We ask for conditions of permission: what structure must a world have for a learner to legitimately receive the framework before the data? The paper gives three interlocking answers, which — together with the falsifiable predictions staked in Section 6 — constitute the paper’s original contributions: 1. The derivation of the four worlds (Section 3) — together with an analytical instrument. From the learner’s point of view, using three variables — data interface, constraint structure, and design rules — we derive the cognitive world as four: the phenomenal world, the basic physical world, the artificial physical world, and the artificial symbolic world. This is not a metaphysical partition but a task-relative one, aimed at the learning task; it explains the historical order of AI’s breakthroughs (why the symbolic world fell first) and predicts where the next breach will occur. Section 8 will use this ruler to take the measure of an independently proposed theoretical construct (R.6) — an instrument that can measure things it was not built to measure is more than a self-portrait. 2. A legitimacy criterion for prior frameworks (Section 4, the heart of the paper). Prior extraction is legitimate if and only if the object world is (i) intentionally constituted — there exists a purposive design, prior to its instances, that constitutes them and defines their norms; and (i) that constitutive process left a readable generative archive. We further argue the criterion’s sufficiency: when both conjuncts hold, the extracted framework binds its instances not by statistical extrapolation but by promulgation, and its failures are alarms rather than silences — which is exactly what “operable intelligence” means in institutional settings. 3. The layering necessity of prior frameworks (Section 5). A prior framework that satisfies the criterion cannot be a flat table of rules: building one pursues four goals — stability, openness, abstraction, concreteness — which pair off, each pair requiring a carrier incompatible with the others; hence the framework is necessarily layered: syntax layer, concept layer, knowledge layer, instance layer. We give a reductio argument for the layering lower bound (≥ 4), the outline of a candidate minimal skeleton, and deployment claims across five industrial domains as an existence proof. A note on positioning. This paper deliberately postpones comparative positioning until after the criterion is built: every comparison in the sequel is performed with our own legitimacy criterion as the measuring instrument, and the instrument must be built before it can measure. The related-work section (Section 8, placed after the objections section) does three things there: it establishes the intellectual lineage of the framework; it draws the boundary between the framework and its nearest neighbors; and it re-arranges existing paradigms along the variable the criterion isolates — whether a knowledge structure is described or promulgated. Accounts with every relevant tradition — knowledge representation and description logics, the semantic web and knowledge graphs, normative multi-agent systems, neuro-symbolic AI, industrial ontology practice, and the agent-engineering lineage of LLMs — are settled there, with physics-informed neural networks (PINNs) as the control group. For the reader who wants a one-sentence answer to “how is this different from existing approaches”: existing traditions engineer explicit structure; we ask when, where, and why such structure is legitimate — and establish that it must be layered. Route Source of knowledge Intelligence before deployment? Requires historical data? Captures normative violation? Cross-instance transfer Data-driven ML observations No High No Low World models dynamics learned from video and interaction Partial High No Medium Knowledge graphs / ontologies human concepts Depends Low Partial Medium Digital twins asset state Partial Medium No Medium Constitutive priors (this paper) design archives Yes Low Yes High Table 1: Routes to physical-world intelligence, compared. Ratings concern physical-world deployment only; whatever normativity a digital twin has resides in its attached rule engine, not in the twin itself. This table is the compressed statement of the framework’s comparative claims: every cell is derived from the analyses of Sections 2–5, and its falsifiable consequences are registered as the predictions of Section 6. This table summarizes theoretical predictions of the framework; their empirical validation is organized as the demonstration register of Appendix C. Section 6 puts the theory on the table: five falsifiable predictions (the continuation of AI’s breakthrough order; the victory map of the prior route in data-scarce design domains; the cross-domain invariance of the concept layer; the ceiling of VLA models entering institutional settings alone; and the irreplaceability of the ontology layer for large language models), with an explicit declaration: if the key predictions fail, the framework itself is wrong — by Lakatos’s standard, the paper accepts the verdict of progressive versus degenerating research programmes. Section 7 restates, in their strongest form, and answers six objections, including the AlphaFold case, the encyclopedia objection (“the LLM has already read every document”), the tacit-knowledge problem, the category mistake of physical-law priors (PINNs), the dismissal “this is just ontology engineering,” and the learning theorist’s challenge, “where is the theorem.” Section 8 settles the accounts of theoretical positioning and related work, with the criterion as instrument. Section 9 ends with a programmatic judgment: the four worlds are all indispensable, but only the designed world can supply AGI with an interrogable skeleton — the flesh is posterior machine learning, the skeleton is prior promulgation; intelligence that can be interrogated is not a luxury but infrastructure. Three methodological declarations. First, our partition and criterion are instrumental: we derive the partition of worlds from the task “how AI learns the world,” and we claim no metaphysical conclusion about “what the world itself is” — in this we belong to the same methodological family as Uexkül’s Umwelt and Dennett’s stance theory (§3.1). Second, all industrial deployment evidence in this paper is given as “claims + commitments” (Appendix C): we state deployment facts and publicly testable invariance commitments, but disclose no proprietary technical detail; the full empirical publication is reserved for a commercially viable moment. Third, scope. The four worlds are derived for machine learners, and the paper’s claims are bounded by what the criterion certifies. We do not propose a path to general intelligence, and we do not claim that any single world is sufficient for it; nor do we propose a theory of open-ended learning in unstructured phenomenal environments, or of biological cognition (§3.2, footnote 3). Where AGI appears (Section 9), it is a programmatic orientation — about which floor of intelligence can be built now, with the tools civilization already ships — not a claim about capabilities, roadmaps, or timelines. The route of the paper: Section 2 re-reads the history of AI from the learner’s point of view and distills three variables — data interface, constraint structure, design rules. Section 3 uses these variables to derive the four worlds, with demarcations from Popper, Hartmann, Simon, and others. Section 4 builds the legitimacy criterion and its sufficiency. Section 5 derives layering necessity and gives the existence proof. Section 6 states the falsifiable predictions. Section 7 restates and answers the objections. Section 8 settles theoretical positioning and related work with the criterion as instrument. Section 9 concludes. 2. Re-reading the History of AI from the Learner’s Point of View 2.1 The Rule Era: Death in a World Without Archives The first death of artificial intelligence deserves a careful autopsy, because its cause has been widely misdiagnosed. Grammar-based NLP tried to teach machines language by first writing grammar as rules and then letting machines parse and generate by them. The rules multiplied; the exceptions multiplied faster; the entire enterprise was eventually crushed under the weight of its own exceptions. Expert systems tried to turn expert knowledge into rule bases; the bottleneck was never the inference engine but the knowledge itself: the rules experts could articulate were always far fewer than what they actually used in judgment. Cyc took this road to its end — roughly a person-century of effort, about a million hand-entered axioms, with millions more cached by system inference [1] — and it proved one thing, though not the one it intended: common sense is not impossible to write down; it is impossible to finish writing down. The common cause of death on the autopsy report is not “the rules were wrong” but the object world had no design rules prior to its instances. Language has no issuing authority; common sense has no design documents — what the rule-writer can reach is forever only the surface regularities of the world, and so hand-written rules are doomed to be descriptions of descriptions. This medical history later re-enacted itself, in another key, in another field: the case of protein structure prediction (discussed in §4.3). The statistical era summarized the lesson of the rule era as “priors are useless” — a misdiagnosis. The correct diagnosis: in a world without design rules, hand-writing “what the rules are” is usurpation. The criterion form of this diagnosis must wait until Section 4, but the medical record is here. 2.2 The Statistical Era: The Prior Changed Its Form, Not Its Existence The victory of deep learning has been narrated as the victory of “abandoning priors”; the narrative is only half true. What the statistical era actually abandoned was hand-written rule content. Two things commonly conflated happened alongside. First, inductive bias was engraved into architecture — convolutional networks carve translation invariance into connectivity, attention mechanisms carve modes of combination into the computation graph; the prior did not die, it remarried, from “rules” to “architecture.” Second — and this is another matter entirely — the archive of the generative process was ingested as data: the multiple sequence alignments of AlphaFold 2 are the most brilliant example (§4.3). There the archive is input, not norm; eating it as data is the standard way posterior learning wins on a happenstance-readable archive, and it has nothing to do with “the prior framework has returned.” Two accounts can now be settled. The statistical era abandoned only hand-written rules — to say “the prior is dead” overestimates what it abandoned. Architectural bias and archive-as-data are not the return of the framework — to say “the prior has resurrected in disguise” underestimates what Section 4 will restore: the promulgated norm itself. The order of the statistical era’s breakthroughs also hides information: board games (fully closed rules) fell first [2]; in the large-model era, both code and natural-language systems entered commercial deployment — code among the earliest, represented by Codex-powered GitHub Copilot [3; 4] — then latent-space diffusion models drove the rapid maturation of high-resolution image generation [5], while robotics lags to this day. The order is not random; it correlates strongly with two variables: a world’s constraint strength (how small the semantic space is) and the quality of its data interface (in what form, at what cost, data is available). The position of code is especially worth marking: it is the only domain whose training corpus is pre-governed by design rules — code that fails syntax and type checking never enters the repository — and whose correctness can be adjudicated mechanically by compilers and tests; the prior is not in the model, but upstream of the data. The success of text has a further, often overlooked invisible condition: language data is written to be read — humanity spent millennia pre-processing language into a free tokenizer: discrete, standardized, low-noise. The physical world offers no such free lunch. Finally, an inverted reading of Moravec’s paradox is worth recording: for humans, sensorimotor skills are easy and symbolic reasoning hard; for AI, exactly the reverse [6]. The cause lies in evolutionary history: sensorimotor competence is a deep capability optimized over hundreds of millions of years, while abstract thought is a shallow, recent invention — and the shallow is easy to learn. This inverted reading will receive its coordinate system in Section 3: AI’s breakthrough order advances along the constraint axis, from strong to weak. 2.3 The Present Predicament: Three Cracks and a Temptation Having been crowned in the symbolic world, the statistical era now shows three cracks on the front where it pushes into the physical world. Symbol grounding remains unsolved. Viewed through the symbol grounding problem [7], the “meaning” of a large language model is a network of relations between symbols, not a relation between symbols and the world. It can talk about water, but it has never touched water; its physical common sense is second-hand, inherited from humanity’s linguistic compression. The thirst for physical data. Physical instance data is scarce, expensive, and privatized: there is no free corpus of plumbing and power, and failure data is rarest of all (the cold-start deadlock of Section 1). The barren lands of data happen to be the rich mines of value — the asset-intensive artificial physical world. The directional confusion of world models. Two main routes currently push toward the physical world. One is physical simulation: simulators such as MuJoCo and NVIDIA Isaac are the main training grounds of robot reinforcement learning — but the precise demarcation is that what they simulate is in fact the artificial physical world (rigid bodies, joints, motors), while the sim2real gap opens mainly at the phenomenal layer (textures, friction details, sensor noise). The other is the learned world model: predictive representations of world dynamics learned from video and interaction data (Ha & Schmidhuber’s World Models [8], LeCun’s JEPA line [9], Genie [10], NVIDIA Cosmos [11]) — a research programme with real ambition and real achievements: predictive dynamics, embodied interaction, scalable simulation. Our demarcation concerns not capability but object of study: a world model learns how states evolve — P(st+1∣st,a)P(s_t+1 s_t,a); institutional intelligence must know what defines a correct state — which states count as normal, which as violations. A world model can be entirely right about what happens next while remaining silent about whether it ought to happen: as Section 1 put it, prediction is by construction description; correctness is promulgated. These are two different problems — and the second one is ours (Section 4). The source of capability is not the source of normativity. One clarification deserves to be made explicit, lest the demarcation be over-read. This paper constrains the source of norms — norms must be promulgated, not induced from data; it has never constrained, and does not constrain, the source of capabilities: a carrier’s descriptive abilities — language, trajectory prediction, perception — may be acquired through any channel, whether the corpus is symbolic text, operational data from real facilities, or the output of an offline physics simulator. The distinction is robust because every output of a carrier remains adjudicable against promulgated norms (⊧ ): the pedigree of a capability is irrelevant to the institution’s trust in it — trust comes not from the corpus but from the tribunal. A judge’s experience is, by curriculum vitae, posterior; the legality of the verdict is, by statute, promulgated. Learning the path beneath promulgated norms is not a concession to posterior doctrine; it is the flywheel itself (§1, §4.4): promulgation supplies the “good”; learning supplies the “way.” Alongside the three cracks stands a great temptation, which must be dealt with before entering the main text: “Large language models have already read everything humanity has written — manuals, specifications, the text of drawings are all in the corpus — so the discussion of prior frameworks is redundant.” The temptation conflates two things: the symbolic shadow of a world and a constitutive framework over the world. An LLM digests constitutive documents as descriptive statistics — it reads a specification the same way it reads a novel — and so it can recite norms, but has no mechanism by which it could be constrained by them; its knowledge is an average of types, without anchoring to instances (it knows what chiller manuals generally say, not whether this chiller at this moment deviates from its own specification); its output is plausible text, not an auditable state. A shadow is not a framework — this distinction is the ground on which the “LLM encyclopedia” objection is fully dismantled in Section 7; here we only establish it, without unfolding it. We also establish here the paper’s position statement: prior frameworks and large language models are not competitors — the framework is a semantic anchorage, giving everything the LLM reads a constitutive, auditable, instance-bound place to land; it is the LLM’s complement, not its rival. The falsifiability of this complementarity structure is registered as prediction P5 in Section 6. 2.4 The Lens: Three Variables Three eras and three cracks can be gathered into one lens. Any world, as presented to a learner, is characterized by three variables: 1. Data interface: in what form, at what cost, does the world present data to the learner? A continuous high-dimensional sensory stream, or discrete standardized symbols? Free tokens prepared by others, or a tokenization one must invent oneself? 2. Constraint structure: how compressible are the world’s internal regularities? How large is the semantic space — closed like chess rules, or open like the weather? 3. Design rules: do there exist rules prior to instances that constitute and regulate them? Are they readable, and for whom were they written? The rule era died of a misjudgment of the third variable (hand-writing rules where no design rules exist); the statistical era’s breakthrough order was set by the first two; the three cracks of the present moment are the hostile values of all three variables in the physical world. The next section uses these three variables to do one systematic thing: derive the complete partition of the cognitive world — four cells, no more, no fewer. 3. Deriving the Four Worlds: Partitioning the World from the Learner’s Point of View 3.1 Methodological Declaration: A Task-Relative Partition, Not a Metaphysics Partitioning the world is a dangerous move, so let us first say clearly what we are not doing. This paper does not answer “what is the world itself made of” — that is the question of metaphysics, and we return it to metaphysics. We answer a much narrower question: for an artificial intelligence that must learn the world, which differences among worlds matter? This way of asking has an explicit methodological lineage. Uexkül’s Umwelt tells us that every species lives in an environment carved out by its own sensory and motor equipment — the tick’s world contains only temperature, butyric acid, and falling; the rest of existence simply does not exist for it [12]: the partition of the world depends on the observer’s equipment and task. Dennett’s stance theory goes further: an object’s “properties” switch with the stance you adopt — the same thermostat is metal and current from the physical stance, a temperature-controlling device from the design stance, and “an agent who believes the room is too cold” from the intentional stance [13]. The ontology engineering of Gruber and Guarino demotes (or promotes) “ontology” from metaphysics to an engineering artifact: an ontology is a specification of a conceptualization, its value adjudicated by the task it serves [14; 15]. The legitimacy of the four worlds is therefore arbitrated not by metaphysics but by learning theory: does it have explanatory power (can it explain the actual order of AI’s history — the full re-reading is Section 2’s burden; §3.2 gives the synopsis) and predictive power (can it issue falsifiable predictions — Section 6)? This is also a self-imposed limit: to read the four worlds as a claim about “being” is to misread this paper — it has no intention, and no capacity, to answer that kind of question. 3.2 Three Variables, Four Cells Section 2’s learner’s-eye re-reading of AI’s history distilled three variables presented to any learner: data interface (in what form, at what cost, the world presents data), constraint structure (the compressibility of the world’s regularities — the size of the semantic space), and design rules (whether rules prior to instances exist that constitute and regulate them). The four worlds are not enumerated; they are the actually occupied cells among the value combinations of these three variables.44 4 Organisms and life phenomena do not occupy a separate cell: their generative archive (the genome) belongs to the “happenstance-readable” gradient of §4.3 and fails conjunct (i) of the criterion (evolution has no intent); they are therefore grouped under the basic physical world. The costly mining history of that archive is the protein case of §4.3. Note that the third variable will be refined in Section 4 into the two conjuncts of the criterion — the rules must be intentionally constitutive, and the constitutive process must leave a readable record; this section asks only whether rules exist; Section 4 asks whether they are legitimate. World Examples Data interface Constraint structure Design rules Phenomenal world cold weather, soft cloth, cool wind continuous high-dimensional sensory stream; no free tokenization weakest (open) none Basic physical world pushing boxes, tiring when lifting acquired in interaction with nature fundamental laws extremely strong; composite systems effectively open through compounding of variables none — laws are discovered descriptions, not promulgated rules Artificial physical world lamps, air conditioners, buildings, machines dual channel: physical instances + design archives strong (dual constraint of physics and intent) yes (rules promulgated by design archives) Artificial symbolic world text, books, programs discrete tokens, written to be read strongest (closed) split: code carries its own; language does not (§4.3) Definition (artificial physical world). The artificial physical world55 5 “Artificial” is used throughout in Simon’s sense: constituted by design rather than by nature — regardless of the designer’s species [16]. The boundary of this world is drawn by the criterion of Section 4, not by biology; the limiting case of authorless artifacts is treated in Remark 4.3. is the class of systems whose norms exist in readable form prior to themselves — (i) its instances are physical: matter, energy, and processes in spacetime; (i) every instance is constituted by a purposive design that exists prior to it, specifying what the instance is and what counts as its failure (C1); (i) the design is recorded in a readable generative archive — drawings, manuals, operating procedures (C2). What the criterion of Section 4 certifies is precisely this class of systems. Two declarations. First, the four worlds stand in composition, not in parallel: they are not four elective ontologies but four floors of a complete intelligence, none dispensable — the developmental order of the child (first the sensory-phenomenal, then intuitive physics, then artifacts, finally symbols) builds these four floors from the bottom up; together, they form a complete stack for machine intelligence. Second, a coordinate axis runs through the four cells: the more artificial, the stronger the constraint and the more finite the semantic space; the more phenomenal, the closer to an open world. This constraint axis immediately yields explanatory power: AI’s breakthrough order — board games first, code and text in succession, images and video after, robotics still lagging — advances roughly from strong to weak constraint; programming became one of the first applications of large models to enter large-scale commercial deployment, and remains the most thoroughly commercialized, because the symbolic rules of code are more closed than those of natural language (§2.2). It also yields a prediction: the next breach is not the weakest-constraint phenomenal world — the direction the world-model and simulation routes aim at — but the second-strongest, archive-carrying artificial physical world (Section 6, P1). 3.3 The Compression Stack: Cognition Abstracts Upward, Engineering Realizes Downward The four worlds are not four adjacent rooms but a stack, and on the stack two movements run in opposite directions. Cognition moves upward. Experience flows from the lower worlds upward, condensing layer by layer: the phenomenal sinks into the common sense66 6 This floor is sometimes called the natural physical world. We decline the name: “natural” covers the phenomenal world equally, while this framework, following the learner’s data conditions, splits the traditional “natural world” into two floors — what the senses can reach, and what instruments and laws can reach. The name marks the partition, not a symmetry. of basic physics (intuitive physics); basic physics and artifact experience sink into language — language sits at the top of the stack, the shadow of the three worlds below. This structure has a corollary aimed squarely at the present: a large language model trained on pure text has never touched the world, yet can tell second-hand physical common sense — because what it absorbed is the shadow of a shadow. A shadow can talk about the world; it cannot stand in for it. Engineering moves downward. Starting from symbols, design documents are realized as physical artifacts — drawings become buildings, code becomes processes; and once running, the artifact keeps sinking downward — buildings stand on rock and soil, machines age under wear and thermal expansion, indoor heat and light return to human skin and eyes: the downward terminus of engineering is the basic physical and phenomenal worlds themselves. Once both movements have run their full course, the midpoint of the stack reveals its unique position: the artificial physical world is the only bidirectionally permeable floor, existing simultaneously as physical object and as symbolic norm — in the transitional language from Section 3 to Section 4, it is the only “natively bilingual” world, the only world that ships with its own source code. The criterion of Section 4 will show why the legitimate prior framework holds here and only here. The stack of the four worlds, the constraint axis, and the two opposing movements are shown in Figure 1. PhenomenalworldFundamentalphysicsArtificial physicalworldArtificialsymbolic worldcognition abstracts upwardengineering realizes downwardconstraint strength: stronger → Figure 1: The four worlds as a compression stack. Constraint strength rises toward the artificial end; the artificial physical world (shaded) is the only layer traversed by both movements — the only world that ships with its own source code. A dividing line deeper than “artificial/natural” follows. The true divide is not artificial versus natural but carrying design rules versus not — and this line does not run between the four worlds; it cuts through the symbolic world. On one side of the line, the artificial physical world (via design archives) and the formal part of the symbolic world share one property: rules precede instances. Code carries rules by definition — the program exists first, the process runs after, and correctness can be adjudicated mechanically by compilers and tests; that programming became one of the first large-model applications to enter large-scale commercial deployment is for exactly this reason (§3.2). On the other side, the phenomenal world, the basic physical world, and language all lack rules prior to instances — language has no rules prior to usage; the death sentence of rule-based NLP was executed long ago (§4.3). The rule-carrying side permits the learning direction “framework first, instances after”; the other side permits only bottom-up. One misreading must be prevented here: permission is not possession — today’s programming models mostly do not explicitly carry a prior framework, yet their success depends precisely on the rule-first nature of the code world: the training corpus is pre-governed by language specifications; erroneous code never enters the repository; the prior is not in the model, but upstream of the data. Section 4 is the full argument of this dividing line. 3.4 Demarcations and Homages The idea of the four worlds has several close relatives; we pay homage and draw boundaries to each — boundary and homage matter equally, because together they fix the paper’s academic coordinates. Popper’s three worlds. Popper divided existence into the physical world (W1), the world of subjective experience (W2), and the world of objective knowledge (W3) [17]. An aircraft is thereby sorted into separate categories: the airframe belongs to W1, the design to W3 — as a unified whole constituted by design, the artifact has no independent place in Popper’s trichotomy. Our framework gives the artifact a home: the artificial physical world is the institutionalized interface of W1 and W3 — here, knowledge not only is “about” the object but “constitutes” it. This interface is exactly the scope of Section 4’s criterion. Hartmann’s emergent strata. Nicolai Hartmann’s stratified ontology divides being into emergent layers — matter, life, psyche, spirit [18]. We salute the stratified intuition but demarcate the goal: his strata are strata of being; ours are strata of learning — the criterion of partition is not the object’s ontological status but the data conditions the learner faces. Simon’s Sciences of the Artificial. The closest precursor: Simon proposed a science organized around the artifact [16]. But he asked “how artifacts are possible and how they are designed”; we ask “how artifacts are learnable” — one might say this paper adds the learning-theory chapter to the sciences of the artificial. Wang Fei-Yue’s parallel worlds. Wang’s parallel-systems approach builds artificial systems corresponding to real complex systems and supports their management and control through computational experiments and real-virtual interaction [19]. What is needed here first is not demarcation but rectification of names: his “artificial world” and our “artificial physical world” share the literal words but not the referent — his artificial world is an artificial virtual world, a digital sandbox for computational experiments; ours is the artificial real world, reinforced concrete plus archives written to be read. The two are therefore not opposed views but opposed semantics: contrasting approaches, different objects. He moves the real into the virtual; we inject the normativity sedimented in the symbolic world into the cognition of real artifacts. Of the two “artificial worlds,” one is virtual, the other is reinforced concrete — and the place where robots will eventually clock in for work is the latter. Kant and the place of metaphysics. The word “prior” must first be cut apart from Kant (see the terminology footnote in Section 1): Kant’s a priori is the legislation the subject promulgates to experience, needing no archive because built into the knower; our prior framework is world-side — promulgated by intentional agents, documented, extractable. The rationalism–empiricism dispute presupposed the dilemma “the prior is either innate or nothing”; this paper points to a third road: the prior can be externalized, documented, and extractable. Looking back at the history of metaphysics from here yields a slightly cruel diagnosis: classical ontology tried to hand-write the content-priors of “being” — an archiveless world; being has no purpose and no readable archive, and so, just as grammatical rules were crushed by the exceptions of the corpus, ontological systems were worn down for two thousand years by counterexamples and schisms; the problem was never the ontologists’ brilliance but that the mine they worked has no ore. The criterion also indicates, in return, the only readable mining district for ontology: classical ontology searched for “the source code of being”; among the four worlds, only the artificial physical world actually ships with one. A readable ontology exists only in the designed world — only what it reads is no longer “being” itself, but the constitutive norms of artifacts. Remark (3.1 — No floor can stand in for another). The four worlds are not four perspectives on one intelligence but its four irreducible compartments. Consider a thought experiment. A person imprisoned in 1996 retains, thirty years later, intact access to two of the four worlds: the phenomenal world still presents itself to his senses, and the symbolic world was never interrupted — throughout those thirty years he could read books and newspapers every day. Yet upon release, he cannot live. The deficit is not in those two intact channels; it is in the artificial physical world: payment systems, turnstiles, smartphones, dispatch and access-control facilities — a floor of constituted artifacts whose norms changed during his absence and cannot be read into cognition through the symbolic channel. They must be re-encountered on their own terms: artifacts that execute their own rules. No amount of reading about them substitutes for operating among them. The intelligence of each world must be acquired in that world’s own way; no floor borrows from another. 4. A Legitimacy Criterion for Prior Frameworks: When May the Framework Be Erected First, and Learning Done After? 4.1 Three Asymmetries Begin with three facts about “which comes first: the symbolic structure, or the world it describes.” Grammar comes after language. Natural language served long before the first grammar was written; and every written grammar is a posterior compression of existing usage — a statistical summary of what speakers were already doing, not the generative cause of those doings. This is the uncontested core of usage-based linguistics [20; 21], and it has an engineering corollary for which rule-based NLP paid full tuition: a rule system distilled from a corpus cannot be turned around to serve as the machine that produces the corpus. The failure of rule-based NLP was not insufficient effort or talent; it was mistaking the descriptive layer of language for its generative layer. Nor did statistical methods succeed because they found better rules — they succeeded because they abandoned the presupposition that “rules live upstream of data” at all. Laws come after phenomena. One floor down, the same asymmetry holds. Newton’s laws stand to the motion of bodies exactly as grammar stands to speech: a consummate, belated compression — the compressed regularity had been running for eons before it was stated. Physics is the most successful descriptive framework humanity has ever built, but physical laws are discovered, not promulgated. The phenomenal world does not consult its own laws; it is the laws that read the world. A learner who wants the framework before the data can therefore only reverse-engineer it from data --- which is exactly what machine learning does at scale, and what no hand-written rules ever did.77 7 Writing physical laws explicitly and injecting them into a neural network’s loss function (as in physics-informed neural networks, PINNs [22]) does not change this asymmetry: the injected laws remain products of posterior compression — humanity’s statistical summary of centuries of observation of the basic physical world, pre-positioned as a training constraint, not norms promulgated to the world before its birth. Their epistemic identity is the same as grammar’s: descriptive rules, posterior to phenomena. §4.3 gives the legitimacy boundary of such descriptive priors. Drawings come before artifacts. Now look at a building, an aircraft, a power station, an air conditioner. Here the symbolic structure — drawings, specifications, bills of materials, control logic, industry standards — is not distilled from the artifact after the fact; it is promulgated prior to the artifact: it regulates the artifact’s generation — what to build, how to build it, what counts as built to standard. A drawing does not describe what a building is; it prescribes what a building shall be. Burn all the drawings, and the construction industry would not become “harder to describe” — it would cease to exist in its present form. This third case is not an edge variant of the first two; it is a different causal topology, and it is the load-bearing observation of this paper. (What exactly the difference is, we hold in place for now — it will become the keystone of §4.4.) The question of this section follows directly: under what conditions is it legitimate for a learner — in this paper’s context, a machine-learning system — to first extract a prior framework and then let the framework govern subsequent learning, rather than being confined to recovering the framework from data after the fact? §4.2 gives the criterion, §4.3 unfolds its gradient, §4.4 supplies its sufficiency. 4.2 The Criterion A natural initial formulation — the one from which this theory in fact set out — is temporal: wherever there is design first, birth after, a prior framework is legitimate. By this criterion artifacts qualify and language does not. But it admits an object that should not qualify: the organism. Genes precede proteins; evolution precedes every individual — in mere temporal order, genes stand to the body exactly as drawings stand to the building. The temporal criterion takes back the very case it meant to exclude, which proves that temporal order is the symptom, not the discriminating variable. (Section 7 returns to this self-criticism; it matters here because the corrected formulation does real work below.) The discriminating variable is not when the symbolic structure exists but what it does. A drawing is promulgated with a purpose: it is written down in order to regulate what this artifact shall be. In so doing it accomplishes what no gene and no grammar can — it constitutes its object and defines the object’s norms. The difference is visible in a linguistic fact: deviation from a drawing is called a defect; deviation from a gene is called a variation. There is no error in evolution, only outcomes; the word “broken” exists only in the artificial world. Purpose, function, failure — the entire normative vocabulary on which engineering, operations, and maintenance run — is defined relative to intentional design, relative to norms and constitutive rules that can be satisfied or violated [23; 24, on constitutive rules; 13, on the design stance]. We call them constitutive norms: the design precedes the artifact not only in time but in reason — it defines what the artifact ought to be. And it is this same difference that separates a prior framework from any “model that has read every human document”: a language model trained on a massive corpus can recite the norms, but has no mechanism by which it could be constrained by them — recitation is not compliance. Constitutive normativity, however, is an ontological property, and a criterion meant to guide learners cannot be purely ontological. A framework can exist yet be out of reach. Stradivari violins and Roman concrete were indisputably intentionally constituted; the knowledge that constituted them is lost, and however definite the purpose once was, nothing can be extracted today. The ontological criterion (“is it intentionally constituted?”) answers what is true; a learning-theoretic criterion must also answer what is accessible. Since the stance of this research is explicitly instrumental — we ask the question to give machine-learning research directions, not to catalogue being — the criterion is forced to carry an epistemic hemisphere. The two hemispheres join into one conjunction: The Legitimacy Criterion. Extracting a prior framework from a world W is legitimate if and only if: (i) W is intentionally constituted88 8 The dual-nature programme of Kroes and Meijers [23; 72] has established that technical artifacts are constituted jointly by their physical structure and their intentional description — an artifact is, in part, what it is for. We inherit this result; our question begins where theirs ends. The dual-nature programme asks what an artifact is; we ask what a learner may extract from the world the artifact inhabits. Their constitution is metaphysical and designer-centered; ours must additionally be epistemic and learner-centered — whence conjunct (i): constitution that left no readable archive (Stradivari violins, Roman concrete) is real but inaccessible, and for a learner, inaccessible is indistinguishable from absent. — there exists a purposive design, prior to its instances, that constitutes them and defines their norms; and (i) the constitutive process left a readable generative archive. Each conjunct kills a class of counterexamples the other cannot. What the criterion certifies is exactly the class of systems defined in §3.2 — the artificial physical world, whose norms exist in readable form prior to itself. Worlds “designed but whose archives are lost” (violins, concrete, every artifact whose archive burned) die on (i): the prior exists in principle but cannot be had in practice. Worlds “readable but purposeless” (the genome, multiple sequence alignments, the fossil record) die on (i): the archive can be read, at no small cost, but what is read out is structure without normativity — which is why, as §4.3 will argue, AlphaFold can recover a protein’s shape from the evolutionary record but can never recover what the protein is for, or what would count as its failure. Normativity travels only with intent. The decision structure of the criterion is shown in Figure 2. A world W, candidate for prior extraction(i) Intentionally constituted?do promulgated norms precede the instances?(i) Readable generative archive?can a learner actually extract it?structure without normativitybottom-up learning onlygenomes, MSAs, the fossil recordprior in principle,unreachable in practiceStradivari violins, Roman concreteLegitimate: framework first, learning afterthe artificial physical worldyesnonoyes Figure 2: The legitimacy criterion as a decision structure: two conjuncts, each killing a class of counterexample the other cannot. Only a world that is both intentionally constituted and readably archived admits a legitimate prior framework — framework first, learning after. One dividend of the conjunction deserves to be pointed out now. Conjunct (i) supplies philosophical completeness: it explains why the priors of a designed world have authority over their instances, not mere correlation with them. Conjunct (i) supplies the practical boundary: it confines legitimate prior extraction to artifacts that are “archived and institutionally maintained” — which happens, not by coincidence, to be exactly the class of artifacts on which operations, maintenance, and engineering value rest. The theoretical boundary and the deployable application boundary coincide, because both are drawn by the same two conditions. Degrees of constitution. Conjunct (i) also admits degrees. The binding force of design is strongest in the artificial physical world — drawings constrain matter that cannot read them. When the regulated object is itself an intentional actor, constitution is “leaky”: chartered organizations (companies, armies, well-defined government departments) are indeed intentionally instituted, and their normativity is real (violating an operating procedure brings discipline, not variation) — but the members continuously interpret, negotiate, and amend the charter’s effective force; an organization is not an artifact “legislated once” but a collective re-legislating itself daily [24, on institutional facts]. Constitutional strength rises with the degree of codification; the military and intelligence agencies are the strong-constitution pole of the organizational world. The criterion’s domain of claim is therefore not the category “physical” but the property “constitutive”: from drawings to statutes runs a continuum — armies and governments fall near the artificial-physical pole (highly codified statutes, highly definite missions), large commercial organizations in the middle (written archives split from living ones), and the grayest civil organizations slide toward the near-ruleless, archiveless end. Our focus on the strong pole is methodological polarization, not blindness: the middle band, where rule-grayness and archive-grayness compound, is better served by sociological treatment, while scientific treatment must first state the strong pole clearly. Once the spectrum is established, every position in the middle band can be located by the two variables of constitutional strength and archive readability — focus does not diminish explanatory power. How this grading superposes on the archive gradient, see §4.3; its corollary for “whether a cross-organization universal framework can exist,” see the organizational-domain analysis in Section 6. 4.3 The Readability Gradient of Archives Neither conjunct of the criterion is binary. The end of the last section explained that (i) admits degrees of constitutional strength; this section unfolds (i): archives differ enormously in whether they are written to be read, and extraction cost scales accordingly; finally we treat a class of boundary objects whose archive readability is split. Written to be read. Engineering drawings, BIM models, CAD programs, data sheets, control-logic books, codes and standards are symbolic artifacts produced in order to carry constitutive information across people and time. They are written for readers. Their semantics are institutionally standardized (naming rules, symbol libraries, classification systems), and their intended reader is precisely the stranger to the project. For a learner this is the cheapest data interface: the world has pre-tokenized itself. Happenstance-readable. The genome was written by no one for anyone, yet it is an archive: a high-fidelity trace of a generative process, preserved and partially decipherable. Reading it is possible but expensive — structural biology paid decades for it, and the reading that finally succeeded is deeply instructive. Both generations of AlphaFold treated the evolutionary record itself — the multiple sequence alignment, a compressed archive of “what evolution tried and what it kept” — as data: the first used neural networks to predict inter-residue distance distributions and turned them into learned potentials, taking a clear lead at CASP13 [25]; the second combined the same record with a substantially redesigned end-to-end architecture [26]. In both systems, the evolutionary archive is an input to posterior learning, not a promulgated norm. Note what the protein case does and does not establish. It establishes that conjunct (i) can be satisfied by happenstance — an unauthored archive can still be mined. It does not establish that normativity is accessible — the MSA records “what survived,” never “what was intended,” and so the recovered framework can rank structures but cannot say what a protein is for or what counts as its failure. The protein is the crucial experiment between the temporal criterion and this one. “Design first, birth after” predicts that prior frameworks should fail outright in evolution’s domain; the archive criterion predicts extraction that is expensive and non-normative. What happened is the latter. No archive. Language leaves usage but no generative archive: there is no document from which sentences are issued. Here conjunct (i) fails entirely, and history’s verdict is correspondingly absolute — no scheme for extracting linguistic priors, however ingenious, survived contact with statistical methods (see §4.4). The legitimacy boundary of descriptive priors. The basic physical world likewise has no generative archive, but humanity has compressed its regularities into explicit text — the laws of physics. Hence an injection method worth registering separately: writing laws into a neural network’s loss function (PINNs and kin [22]). This does work on simple boundary-value problems, but its legitimacy is unrelated to our criterion: a law is a discovered compression, not a promulgated norm; it is retrospective — drawn from centuries of posterior observation and pre-positioned as a training constraint; it supplies constraint boundaries (what is physically impossible), not normative semantics (what counts as a fault). The scope of descriptive priors is therefore strictly limited to that small set of problems “approximable as closed physical systems with exactly specifiable boundary conditions”; once one enters multi-scale, multi-physics, intent-bounded complex systems (building indoor environments, weather), it meets exactly the silent failure §4.4 will unfold — the physics residual can still be zero while a fault has already been constituted in engineering semantics, because “shall maintain 24∘C” is not a subset of any physical law. In one sentence: the descriptive prior is the archiveless world’s expensive substitute; in a world with archives, it yields to the constitutive norm. The two belong to different epistemic categories and constitute no counterexample to the criterion (Section 7, O5, unfolds this). The middle band: chartered organizations. One class of objects occupies not one position on the gradient but two. A purposeful human organization simultaneously has two archives: a written archive — charters, statutes, procedures, written to be read, at the gradient’s most readable end — and a living archive — the way things actually run, much of it tacit knowledge that cannot be fully articulated or encoded [27] — which, by this paper’s calibration, slides toward the “no archive” end. Reading the written archive is cheap; reading the living archive is expensive — expensive enough to have created a dedicated job title: the forward-deployed engineer, stationed inside the organization, transcribing the living archive into computable semantics through interviews, shadowing, and immersion. The criterion’s prediction here is concrete: wherever the archive is thus split, the cost of building a prior framework must include a “human reader” fee, rising with the share of the living archive. The organizational case also shows the superposition of archive grading and constitutional strength: the more codified the mission, the more the written and living archives coincide, and the higher the value density of a prior framework — which explains why the most successful deployments of such frameworks concentrate in the most codified organizations, defense and intelligence. As for “whether a cross-organization universal framework can be distilled,” the criterion’s structural negative and its boundary conditions are left to the organizational-domain analysis in Section 6. Extraction cost rises along this gradient — we conjecture superlinearly, though this paper does not attempt to formalize the rate. What the argument needs is only the ordering: written to be read ≺ happenstance-readable ≺ no archive — and among the non-symbolic worlds, only the artificial physical world occupies the first cell. The gradient also settles, in passing, an old historical account. The failure of rule-based NLP is often read as the failure of “priors,” but re-read on this section’s scale, the cell in which it failed is obvious: hand-writing content-priors for an archiveless world — those “rules” were descriptions of descriptions, and the enterprise was crushed under the weight of its own exceptions. The engineer writing control specifications for a chiller plant stands in exactly the opposite position: she is not describing a regularity she observed but promulgating one her colleagues will implement, and her document is archived, standardized, and maintained by the institutions around her. Erecting content-priors over the artificial physical world is not a re-enactment of NLP’s mistake; it is the one situation in which that mistake does not apply — because both conjuncts hold here, simultaneously and cheaply. This is also the precise meaning of the artificial physical world’s uniqueness among the non-symbolic worlds (see Section 3): it is the only physical world that ships with its own source code — purposive, normative, written to be read; every other world must be learned bottom-up from instances, and only the designed world permits the other direction: framework first, instances after. (The life-and-death records of the various injection methods for priors — hand-written rules, dissolved into data, engraved into architecture — are the topic that opens Section 5.) A guardrail against misreading is needed here. The “purpose” of this paper is institutional through and through — purposes are written in requirements documents, functions in specifications, failures in failure-mode analyses; it has nothing to do with any metaphysics of “nature’s designer” [13: the design stance is a strategy, not a theology]. Where institutions stop keeping such archives, the criterion stops applying, and the framework honestly exits its own territory (see §5.5, Section 7). Self-stated scope boundary — the framework applies to artifacts of “high utility value, institutionally maintained”: only such artifacts merit humanity’s full sets of procedures and specifications (the practical boundary of the §4 criterion); the consumer-goods domain (if broken, buy a new one) is outside the domain of claim. The organizational domain is located on the same continuum: codified organizations (defense, intelligence, emergency response) have the highest value density, but the “human reader” cost is structural (§4.3); this paper’s deployment claims and wagers are confined to the physical domain, and the structural corollaries for the organizational domain are in the appendix note of Section 6. A further corollary from §5.3: since failure types are enumerable, an event “wholly outside the known type list and not interpretable as a superposition of known types” can thereby be defined as force majeure — as war and natural disaster stand to operations services; the framework does not fail on such events, just as an insurance contract does not fail on war. The framework claims no victory beyond its boundary — this is the same guardrail as §4.3’s: where institutions stop keeping archives, the framework honestly withdraws. 4.4 Why the Criterion Suffices: The Inversion of Direction of Fit The criterion has so far been tested by counterexamples: each conjunct kills a class the other cannot. But that answers only why it fails without the two conditions, not yet why it works with them — why a prior framework satisfying the criterion can build operable intelligence rather than a bundle of legitimate but useless statements. Answering requires the keystone of the whole section: direction of fit. Borrowing the distinction of Anscombe and Searle [28; 29], we distinguish two directions of fit between a rule system and its object. The direction of a descriptive rule is rule-fits-world: if a grammar has exceptions, the grammar is wrong; if a law is falsified, the law is revised. The direction of a constitutive rule is world-fits-rule: if a building deviates from the drawings, the drawings are not wrong — the building is; the party at fault is called a “construction defect,” or a “fault.” The three asymmetries of §4.1 are, at bottom, exactly this difference: grammar and law stand in the first direction; drawings in the second. Why this difference bears weight can now be stated: descriptive rules have no binding force over the world, and any “framework” extracted from them is a post-hoc ratification of the world, forever carrying inductive risk — the next instance may always be a counterexample; constitutive rules have promulgative force over the world, and a framework extracted from them inherits that promulgative force itself. With this keystone set into the criterion, sufficiency stands on three legs, each corresponding to one dimension of “operable.” First leg: accessibility supplies the head start. The archive is readable before the instances; the framework can be fully extracted before encountering any instance — operable intelligence holds from the first day of deployment, without waiting for posterior data to accumulate. The cold-start problem is thereby untied. Second leg: constitution supplies induction-free generalization within the constitutive domain. A statement extracted from a constitutive archive binds future instances not by statistical extrapolation but by constitutive constraint: instances are made to satisfy the framework; the framework is not guessed right. A statistical model’s validity on new instances is probabilistic — trustworthy in-distribution, mute outside it; a prior framework’s validity on new instances comes from promulgation — so long as the object remains the artifact built, accepted, and maintained according to its design, the framework’s statements about it continue to hold, on the strength not of sample size but of the norm itself. Third leg: normativity supplies alarm-style failure. Instances will of course deviate from design — wear, misuse, aging, drift of operating conditions. A statistical model meeting an out-of-distribution sample fails silently: it outputs as usual, only no longer credibly; a constitutive framework meeting deviation alarms with an address: the deviation is classified on the spot by the framework’s normative semantics as trend, symptom, or fault. In the operations context, “operable” means precisely not never failing but never failing unreported — an intelligence that does not know it is wrong is unusable in institutional settings. The framework’s semantics has its own failure modes built in; this is what posterior learning cannot learn from data, because the word “wrong” is not in the data — it is only in the norms. If the criterion holds, its practical consequence is already visible: in design domains with readable archives, operable intelligence need not wait for large-scale posterior data. What shape this archive should take — history’s three deaths having vetoed the flat rule table — why it must be the four-layer structure of syntax, concept, knowledge, and instance data, is Section 5’s business. Remark (4.1 — Why not simply learn the norms?). One predictable objection remains: since operational data of the artificial physical world will eventually accumulate sufficiently, why not simply learn the regularities and skip the archives? Two mutually independent reasons answer it, both structural rather than operational. First, instability. A statistically induced “rule” is a summary of a data distribution and drifts with it; a constitutive norm is defined precisely by its independence of any distribution. A learned copy is not merely an imperfect approximation of the norm — it does not know it is a norm. It carries no violation semantics: deviation from a statistical regularity is indistinguishable from noise, whereas deviation from a promulgated norm is an event — with direction of fit, a responsible party, and a required disposition. Second, uneconomy. Spending data, compute, and calibration to re-derive an already promulgated, freely readable document buys an inferior and unaccountable copy — while the original sits on the shelf, in clear handwriting. To learn that “goals scored equal goals conceded,” one need not run a longitudinal study of Serie A — the rulebook promulgated it; as an old piece of sports-commentary humor has it: “we are amazed to discover that in this round of the league, goals scored and goals conceded are exactly equal.” Posterior learning is indispensable — it is the flesh. But it must not be allowed to counterfeit the skeleton. This is also the deep content of prediction P4: the decay of the data-first-mover advantage is not data depreciating, but data being unable to buy what only promulgation can supply. Remark (4.2 — Constituted but unread: the regularization channel). It matters that the criterion’s two conjuncts are of different kinds. Intentional constitution is an ontological condition: it asks whether norms exist that in fact constitute the instances. Readability is an epistemic condition: it asks whether those norms can be legitimately obtained. The two can be separated clause by clause. A design convention may genuinely govern the construction of an artifact — deviation from it gets corrected, not merely noticed — yet never be written down. Such a norm satisfies the ontological condition but not the epistemic one. The criterion is therefore distributed: the archive is not a bulk property of a world but a ledger registered clause by clause; every clause of a framework carries its own epistemic fineness. Three design rules follow. First, detection: constitutive-but-tacit norms are detected by the sanction test — violating a descriptive convention invites surprise; violating a constitutive norm invites correction; even without documentation, the direction of fit remains observable in a community’s enforcement behavior. Second, quarantine: a reconstructed norm enters the knowledge layer — never the concept layer, because such norms are concrete and bound to particular systems — wearing an unratified tag: it may participate in reasoning, may be overturned, and is downgraded in alarm attribution. Third, ratification: promulgation is an act, not an epistemic property. Consulting accountable experts, drafting the clause, obtaining an authority’s sign-off, filing it into project documents — this sequence manufactures the missing archive; at that moment the clause crosses the waterline, from flesh to bone. Archive completion is thus itself an independent engineering activity, and conjunct (i) of the criterion should be read as covering “archives that can be produced,” not only archives that happen to exist. What does not follow is permission to bypass the ledger: inducing norms from behavior without ratification is precisely the induction that Proposition A.5 bounds — and a norm no one has signed is a norm no one answers for. The same machinery serves the gray band of §4.2: organizational norms have a disproportionate tacit share, and the sanction test, the unratified tag, and the ratification channel are the standard operating procedures for working on the gradient — rather than pretending the gradient is binary. Remark (4.3 — The authorless artifact). A thought experiment can sharpen the criterion’s edge. Suppose a machine is designed entirely by an AI system and 3D-printed: no human intent, no readable documentation. Anyone sees at a glance that it is an electric fan; in use, it is one too. When it stops turning, we say without hesitation that it is broken — but note where this “ought” comes from. The artifact itself carries no promulgated norms; “broken” is judged against our archive of fans — airflow, duty cycle, temperature-rise limits, all promulgated by human engineering practice. The object is readable not because it ships with norms, but because it landed in a world that ships with norms. Its readability is borrowed — and this borrowing is precisely the criterion at work. Reverse-engineering the artifact is of course possible — but the product of reverse engineering is hypothesis, not norm. Such reconstructed rules may enter the knowledge layer under the unratified tag of Remark 4.2; and here even the sanction test is unavailable: no community will come to correct a misreading. If one used the reconstructed framework to interpret the AI’s next generation of designs, one would run head-on into the direction-of-fit knot: when the next generation deviates, you cannot tell whether the AI violated its design norms or our reverse engineering guessed wrong. The framework is left with only the world→ direction — every deviation revises the model — which is exactly the Gold bound of Proposition A.5, live on stage. The deeper clarification: the criterion is not anthropocentric. This artifact is excluded not because an AI made it, but because no agent — human or AI — promulgated its norms into a readable archive. Promulgation is an act, and any signer counts: a design system that signs, finalizes, and publishes its own rules brings its artifacts within the criterion’s scope. The boundary of the artificial physical world is drawn by promulgation, not by species — where someone signs, the artifact enters the domain; anyone may sign. 5. Layering Necessity: What a Legitimate Prior Framework Must Look Like 5.1 Lemma One: Design Comes First, the Skeleton Is Ready-Made By the criterion of Section 4: the artificial physical world is intentionally constituted, and its constitution left archives written to be read — both conjuncts satisfied, extraction cost at the lowest end of the gradient. Building a prior framework over this domain is therefore legitimate, and the framework’s skeleton need not be invented: purposes, functions, and norms can be taken directly from the constitutive norms promulgated by the design archives. That purposes exist before artifacts is a fact unique to the artificial physical world. But a skeleton is not a building. The archive tells us where the framework’s content comes from; it does not directly tell us what shape the framework should take — content can be piled into a warehouse or built into a tower. To answer the shape question, one must first ask a prior question: what do we want this framework to achieve? 5.2 Four Construction Goals A prior cognitive framework — the constitutive prior of the title — is an engineering artifact (§3.1), and an engineering artifact’s shape is decided by the specification it must meet. And the specification is not ours to pick: it grows out of the world picture of §3 and the direction-of-fit inversion of §4.4 — exactly four goals, each with its provenance noted in its paragraph. Goal one: stability. The framework’s conceptual core must be constant and must not drift with data. Its source is the second leg of §4.4: the framework’s authority over instances comes from promulgation rather than extrapolation, and promulgation presupposes that the norm itself stands firm. If the conceptual core drifts with data, the word “fault” immediately loses its definition — the norm is the only source of the word “wrong.” A referee who revises his own rules every day is no referee. Goal two: openness. The framework must keep ingesting the new — new instances, new operating conditions, new failure modes — without rewriting itself. Its source is the inherent growth of the real world: humanity’s technological development keeps creating engineering marvels, keeps building new artifacts — the list of objects the framework serves grows forever. A popular misstatement must be corrected here: there is no “symmetric tension” between stability and openness, as if two equal enemies pulled at each other. Stability has no enemy in itself — a dead framework that cares nothing for openness is forever stable. Openness is the threat to stability: it is because new things must still be ingested that stability needs defending. The real question is therefore not “how to compromise between two opposing demands” but: how to grow without drifting. This formulation will do repeated work below. Goal three: abstraction. The framework’s experience must be accumulable in conceptual form — what is learned on one concept applies automatically to all its future instances, rather than each instance being modeled from scratch. Its provenance is the first posture of the act “constraining instances” — for a symbolic structure to constrain instances, it must function as a type reusable across instances. The goal’s source is not data scarcity: even with infinitely rich data, pure instance-eating learning would be pointless — a waste of compute, not a construction of intelligence. Abstraction is the prior framework’s native benefit: it compresses information by nature. A prior framework with poor abstraction loses the very meaning of “prior” — the extreme of non-abstraction is all instances, which is just posterior data learning. Abstraction is the precondition of iteration and accumulation: the same semantics reused across instances is what lets experience snowball. Goal four: concreteness. The framework’s concepts must be able to speak down to concrete physical instances — this chiller, this ward, this pipe segment — judgments have business value only when anchored to instances. Its provenance is the second posture of “constraining instances” — functioning as a clause anchored to concrete objects. Otherwise, however complete the concepts, the framework talks only to itself: that is metaphysics, not engineering. Before unfolding, three old accounts must be settled, because the words stability, openness, abstraction, and concreteness all have prior owners elsewhere. First account: the stability–openness problem has a famous ancestor in machine learning — the stability–plasticity dilemma [30], on which the continual-learning literature still works. But that is same name, different thing: the “stability” that literature protects is learned parameters not being washed out by new data, and “openness” is continuing to eat new data — an engineering problem inside posterior learning, whose typical solution is freezing the lower layers for fine-tuning. Our stability and openness are construction goals of a prior framework, and we have what they lack: norm promulgation as the source of stability (§4.4). Continual learning defends stability by algorithm; ours holds by promulgation. Second account: the abstraction–concreteness layering has a half-century history in knowledge representation — description logics separate the terminological layer (TBox: abstract, design-time, slow-changing) from the assertional layer (ABox: instances, run-time, situation-specific) [31]. We inherit the distinction; §5.3 will show that our layers are divided by goal structure. Third account — homage rather than demarcation, most deserved: that there is an axis between abstraction and concreteness is not this paper’s original claim. Rasmussen, in the process industries (nuclear power), long ago characterized how operators understand the work domain with five abstraction levels — functional purpose → abstract function → generalized function → physical function → physical form [32; 33] — the earliest, and to our industry the closest, empirics of this axis. Our novelty lies elsewhere: first, Rasmussen’s levels are cognitive structures in the operator’s head, reverse-engineered from external observation; what we argue for is a generative structure promulgated down into the design world’s archives — the same axis, a psychological fact there, an ontological fact here. Second, in this paper the axis is not assumed but derived from the construction goals (§5.3) — the division of layers is uniquely determined by the goals. What Rasmussen saw receives here an ontological explanation: it is the cognitive develop-print of the design world’s generative structure. (Another independent layering practice in industry — Palantir’s Ontology — is settled together in §5.5.) Fourth account — the nearest philosophical neighbor on the layering axis: Floridi’s method of levels of abstraction (LoA) [73] establishes that any analysis of a system is relative to a chosen level of abstraction, and that the choice is dictated by the analyst’s purpose. We inherit the purpose-relativity; the demarcation is one of direction. Floridi’s levels are epistemic instruments, chosen by an observer facing a system already built — many LoAs are legitimate, none is ontologically privileged. Our layers are generative facts, promulgated by the design archive before instances exist — the four-layer skeleton is not one analysis among many but the shape the world itself shipped with (§5.4). Where Floridi’s analyst selects a level, our engineer inherits one. The two are complementary readings of the same axis: what purpose selects in analysis, purpose fixed in design. Each of the four goals is necessary, checkable one by one by ablation — remove any one, and what does the framework lose: Remove Consequence stability the conceptual core drifts with data; “fault” loses its definition — the promulgation leg snaps on the spot openness the framework is trapped in staleness: new artifacts, conditions, and failure modes cannot enter; coverage stops at promulgation day abstraction degenerates into an all-instance record — i.e., posterior learning itself; “prior” loses all meaning concreteness concepts talk to themselves, anchored to neither this chiller nor that ward — metaphysics, not engineering With the four goals on the table, the shape question can be posed precisely: should the four goals be realized by one carrier together, or each in its own place? The answer comes in two steps: first see how the four goals pair off and what carrier each pair needs (§5.3); then close with the layering lower-bound proposition (§5.4). 5.3 Four Goal Pairs, Four Carriers The four goals are not four things piled at random. Pairing them along two axes — rate of change (stable × open) and degree of abstraction (abstract × concrete) — yields exactly four goal pairs; and this paper’s claim is: each goal pair needs a carrier, the four carriers are mutually incompatible and cannot be held by one and the same thing — therefore a framework must be layered. The two axes are not chosen: the two poles of rate of change come, one from the definition of promulgation (§4.4 — a norm must precede instances and constrain all that follow), one from the world’s growth (§3 — the object list never closes); the two poles of abstraction come from the two postures of “constraining instances” (type and clause — see goals three and four of §5.2). Each axis has its provenance, and the completeness of the four cells inherits from the axes’ definitional status, needing no separate defense. As for a “fifth goal”: if it truly exists and is incompatible with all four cells, the lower bound automatically rises from four to five and the theorem only gets stronger — this framework bets it will not appear (§5.4 stress-tests this; Section 6, P3, registers the wager). Stable × abstract: the syntax layer — the framework’s foundation. The grammar of the whole framework: the most basic word classes and combination rules. It must be the most stable, nearly unchanging, because it is the expressive medium of every superstructure; and the most abstract, because it denotes nothing concrete. The syntax layer’s constancy is not conservatism but the precondition for the framework to be one system: with word classes and combination rules unmoved, every augmentation of the upper layers forever lands in the same coordinate system. In practice this layer can be astonishingly small — our syntax has only 12 nouns, 47 semantic connections (consolidated into 32 relation terms), plus two levels of adjectives (attributes, and attributes of attributes — the former modify nouns, the latter modify adjectives) — yet it suffices to write the artificial physical world across multiple industries. Open × abstract: the concept layer (also called the semantic layer) — the home of concepts. Purposes, functions, failure-mode types live on this layer. Its semantics are unique and abstract: every concept has a definite meaning and can be collapsed along the hierarchy — the same semantics reused across instances, so experience can accumulate and iterate. It is open: as the range of the artificial physical world the framework serves expands, new concepts are added. But its openness does not threaten stability, because augmentation ≠ drift: the foundation does not move; concepts are only added, never altered. Why can augmentation be add-only? Because this layer’s content is under two convergent constraints: purposes and functions are finite sets — artifacts are made for purposes, a facility’s purpose list (cooling, power, fire protection, egress) is written in requirements documents, and functional-basis research in mechanical design likewise shows that a broad range of product functions can be represented by a standardized and compact vocabulary [34]; failure-mode types are likewise a finite set, because the abstract ways equipment can fail are fixed by materials, energy flows, and causal structure — our engineering practice supplies an empirical claim for the latter convergence: across the facilities served, cross-industry onboarding has never required inventing new failure types, only instantiating known ones. This claim has its everyday form in the engineering inventory: in the framework’s failure-mode (FMEA) knowledge base, physical-system failure mechanisms, medium-state failure mechanisms, system transfer factors, and environmental factors are all maintained as enumerative lists — “types are enumerable” is not a philosophical posture but a table in the warehouse (four lists, 32 classes in all; full text in Appendix B). Stable × concrete: the knowledge layer — codified rule documents. Failure-mode instances, equipment procedures, and operating-condition criteria live here. It is concrete: every rule describes a concrete category of phenomenon or object; and it is stable — but this stability rests not on enumeration but on clause-by-clause codification: the instance set is an open set under continuous augmentation (every new industry, every new device defines new instances), while any knowledge instance, once described and defined, becomes a document, invariant across projects, scenarios, and time. The set is forever open; every clause is forever unchanged. This cell must be carefully distinguished from the previous one — the most easily confused pair of the section: “open” says whether the system can extend to describe more new things; “concrete” says whether it is a concept (the commonality of things) or a concrete phenomenon/object. The concept layer is open and abstract; the knowledge layer is concrete and stable — one governs “what can be said,” the other “what has been said and settled.” Open × concrete: the instance layer — where posterior data lands. This layer holds not the real world itself but the real data describing its instances: continuous telemetry, observation records, streams of operating conditions — posterior data enters the system here. It is open because the real world keeps running, keeps producing new data, never finishes; it is concrete because every datum anchors a concrete instance. This layer is one link of the whole intelligence system’s operation, and the landing ground of the flywheel of §1 and §4.4: the prior starts, running produces data, data flows back — and lands here. It is also a structural humility: the three layers above are all sections of the real world; only this layer presses directly against it — every concept and every item of knowledge must ultimately be cashed out here, or alarm here. Why can the four carriers not moonlight for one another? First nail down “carrier”: a carrier = a content region with a unified revision discipline and a unified denotation discipline. Incompatibility thereby turns from assertion into checkable proposition: rate-of-change incompatibility — the stable pole requires a revision discipline of clause freezing, the open pole requires one of continuous ingestion, and one carrier can have only one revision discipline; denotation incompatibility — the abstract pole requires a denotation discipline of open types (infinitely instantiable), the concrete pole requires one of anchored instances, and one carrier can have only one denotation discipline. Hence: to let syntax be open is to let the foundation grow with the tower; to let semantics be concrete is to lock concepts onto single instances and lose reuse; to let knowledge instances be abstract is for documents to lose their anchored objects; to let the framework touch the open-concrete stream of the real world directly is the old hungry-and-silent road of end-to-end statistics (§1, §4.4). The impossibility of moonlighting is the necessity of layering. The cleverest merger proposal deserves mention: put the four cells in one layer, each entry carrying its own tag. It does not work — a content region whose entries each have their own discipline is already layered; the tags are the layer boundaries; allowing one entry to carry two disciplines amounts to admitting two sub-carriers in the region. The merger changed its name; the layering did not disappear. 5.4 Proposition: A Lower Bound on the Number of Layers The four-layer skeleton, the four goal pairs, and the two families of directed edges (downward instantiation, upward collapse) are shown in Figure 3. Syntax layerstability × abstractionConcept layeropenness × abstractionKnowledge layerstability × concretenessInstance layeropenness × concretenessinstantiation (down)collapse (up)posterior data streams in Figure 3: The four-layer skeleton: one goal-pair per layer; inter-layer relations are of one kind only — instantiation read downward, collapse read upward. The reduction path is the explanation path. By §5.2, a qualified prior framework must achieve four goals; by §5.3, the four goals pair off and each pair needs a mutually incompatible carrier. But before stating the theorem, the connective tissue must be added — otherwise four layers are just four floor slabs. The framework’s structure has two kinds of relations. Within a layer, relations are horizontal: syntax has combinatorial relations with syntax (how word classes collocate), concepts relate to concepts (how they connect), knowledge instances combine with knowledge instances, and instance data has physical-connection relations internally — each layer is its own graph. The combinatorial relation among knowledge instances deserves a sentence of its own: real-world failures often strike as superpositions of multiple atomic patterns, but combination produces no new types — any superposition of element-level failure modes is still a superposition of those types. Combinatorial complexity is therefore a problem of the intelligent-engineering layer (how to decompose and attribute), not an openness problem of the framework layer; the framework’s stability does not collapse under too many combinations. This also cashes out the key promise of interpretability: most of what the field calls “unmeasurable phenomena” and “inexplicable faults” are not new regularities but superpositions of known types — the combinations are complicated, the concepts are old; if the superposition can be translated back into a combination of concepts, the model is interpretable. Complicatedness does not necessarily destroy interpretability — provided there is a set of concepts that can translate the complicated back. The contrast with a purely distributional model is sharp: such a model can report that an input is unlikely; it cannot decide whether the input violates a norm or instantiates a novel but legitimate behavior — for it, “is this a new failure type?” is not a question answered wrongly but a question asked wrongly, and the distinction is not scholastic, because the two cases demand opposite dispositions: one shall be learned, the other must be stopped. Between layers, the relation is vertical, and there is only one kind: abstraction and instantiation (parent–child sets, taxonomy) — knowledge instances instantiate concepts, concepts instantiate syntax, instance data instantiate knowledge instances. Top-down, it is instantiation layer by layer; bottom-up, collapse layer by layer: any concrete phenomenon, however complicated, traveling up the vertical edges, necessarily lands on some stable concept of the concept layer. The reduction path is the explanation path — every judgment gives its reasons up the edges; every failure reports its location up the edges. Therefore: Proposition (layering lower bound: a prior cognitive framework satisfying the criterion has, on the two definitional dimensions enforced by the promulgation structure (time: promulgation precedes and constrains the instance stream; denotation: constraint reaches instances as type or as clause), a layering number of at least four). Argument (reductio). By §5.2, a qualified framework must achieve four goals: stability (S), openness (O), abstraction (A), concreteness (C); by §5.3, the four goals pair along rate of change and degree of abstraction into four goal pairs: S×A, O×A, S×C, O×C, each needing a carrier. Define two carrier incompatibilities: rate-of-change incompatibility — one carrier cannot simultaneously satisfy S (remain invariant across time and instances) and O (continuously absorb new content without rewriting itself); abstraction incompatibility — one carrier cannot simultaneously satisfy A (denote types, reused across instances) and C (anchor single instances). Hence every merger among the four goal pairs triggers at least one incompatibility: (S×A, O×A) and (S×C, O×C) share the abstraction degree, but S conflicts with O; (S×A, S×C) and (O×A, O×C) share the rate of change, but A conflicts with C; the diagonal mergers (S×A, O×C) and (O×A, S×C) conflict doubly (the itemized check of all six mergers is in Appendix A.3). Therefore no two goal pairs can be borne by one carrier. By the pigeonhole principle, if the framework had only three layers, at least two of the four goal pairs would have to crowd into one layer — contradiction; two layers and one are beyond discussion. Hence the number of layers ≥ 4. ∎ One clarification. The proposition establishes only a lower bound and does not claim “exactly four”: an engineering implementation may subdivide a layer (for instance, splitting the knowledge layer into industry sublayers) — an implementation choice that does not touch the theoretical skeleton — the skeleton cannot be compressed further; the flesh can be divided further. (Its falsification channel is registered as Section 6, P3.) Remark (fifth-goal stress test). Once the lower bound holds, the strongest rebuttal is not to question the four cells but to search for a fifth. We test the most plausible candidates on the objector’s behalf: real-timeliness — a performance constraint on carriers, not a new content dimension; safety — that is promulgated content itself, entering the knowledge layer (stable × concrete); interpretability — a byproduct of the layered structure (the reduction path is the explanation path); economy — a matter of the value function, constraining choices outside the framework, not goals inside it. Every candidate reduces to a constraint on the existing two axes, not a new dimension. Further attempts are welcome: produce one irreducible fifth goal, and this proposition upgrades — exactly the kind of wager Section 6, P3 registers. The layering lower bound has a geometry, worth stating last because it welds Sections 3 to 5 into one piece. Look back at the goal map: the artificial physical world is the only world that natively spans the entire map — its generative archives live exactly on the “stable × abstract” side (written to be read, promulgating norms, prior to instances), and its physical instances live exactly at the “open × concrete” corner (continuously running, continuously producing new data, never finished). Every other world must be learned bottom-up from instances; only the designed world leaves the factory already layered — the precise meaning of Section 3’s “ships with its own source code” is here: in this world’s source code, the directory structure of syntax, semantics, and instances is ready-made. The four-layer skeleton the lower-bound proposition requires is exactly the shape this world already has — the framework is the mirror of the world. This is where the criterion (§4) and the proposition (§5) join: the criterion says this world deserves framework-first; the proposition says the framework must have at least four layers; and layering is possible, and cheap, because the world itself leaves the factory pre-layered. Proposition (5.1 — Closure is a condition of artifactuality). The concept layer of the artificial physical world admits a finite closed basis. Argument. The claim is a priori, not enumerative. Suppose some practice introduced a purpose no functional basis can describe. A purpose, as an engineering category, is intelligible only if it can be decomposed into the constitutive dimensions of the artificial physical world — physical effects (transfer of energy, matter, information, or mechanical work) and spatial relations — because these are what a designer can specify, a builder can realize, and a maintainer can inspect. A “purpose” that refuses such decomposition cannot be designed, cannot be built, cannot be maintained: it cannot become an artifact at all, let alone an industry. Closure is therefore not a historical accident but a condition of artifactuality: some finite basis must suffice, otherwise nothing is engineerable at all. The argument guarantees the existence of a closed basis, not that today’s enumeration is that basis; Appendix B’s vocabulary is the current best candidate, and prediction P3 states the public criterion by which its candidacy would be defeated. 5.5 Two Independent Witnesses: A Casualty Report and a Convergence The lower-bound proposition is derived. But it happens to have two witnesses from outside — a natural experiment an industry ran at real cost, and a convergent construction from an independent mathematical tradition. Neither is an illustration of this paper; they are two independent pieces of testimony. Witness one: the merged middle two layers — a casualty report. This section’s lower-bound argument is existential; by itself it does not describe how a compressed framework fails. The automotive safety industry has run this natural experiment for us at full scale, and it supplies the casualty report. Operational Design Domain (ODD) practice [35; 36] requires manufacturers to declare the conditions under which an automated driving system is certified to operate — and in prevailing practice, the attribute taxonomy and the vehicle-by-vehicle declarations are written in the same natural-language document. The concept layer and the knowledge layer are thereby merged into one carrier, and the predicted pathologies break out with clinical regularity: Retroactive reinterpretation. Frozen declarations cite vocabulary entries; once the taxonomy is revised, the meaning of untouched declarations quietly drifts — the certificate’s words have not changed; what they certify has. This is exactly the silent failure mode this framework exists to exclude, and it appears precisely where the two layers are conflated. Cadence deadlock. The taxonomy must stay open — new scenario categories arrive with new technology — while certification declarations must stay frozen. A single carrier cannot keep both cadences: freeze the vocabulary, and the framework is trapped in obsolescence; keep it open, and the certification drifts with revisions. One of the goals must be betrayed; this is the carrier incompatibility of Proposition A.2 instantiated in regulation. Blast radius. With layers separated, a vocabulary revision triggers selective re-instantiation of only the affected declarations; merged, every correction of an ill-defined category voids every declaration that cites it — maintenance becomes total war. Undecidable liability. Liability asks: at the moment of the accident, was the vehicle within its certified domain? Under retroactive drift, the question has no time-stable answer — in-domain by the vocabulary at certification, out-of-domain by the vocabulary at trial. Insurance and adjudication need frozen declarations over a stable vocabulary; the merged carrier pulls out exactly this foundation. The four pathologies reconcile line by line with industry practice: Pathology predicted by the proposition Observation in ODD practice OpenODD’s remedy retroactive reinterpretation (frozen declarations drift with the vocabulary) certificate wording unchanged; certified content changed separation of vocabulary and declarations (ISO 34503 / PAS 1883) cadence deadlock (open vocabulary and frozen certificates incompatible) freeze and go stale, or open and drift binary in/out membership semantics blast radius (one vocabulary fix voids all citing declarations) correcting a mis-defined category empties all citing declarations lifetime-validity clauses for declarations undecidable liability (no stable answer at the time of accident) in-domain at certification, out-of-domain at trial satisfaction semantics (⊧ ) for membership The industry’s symptoms are on record. An industry-wide normative standard was long absent: manufacturers could each only implement mutually incomparable schemes [37]. The ASAM OpenODD standardization project was chartered precisely against these symptoms — its motivation statement names the ambiguity, non-exchangeability, and machine-unprocessability of natural-language ODD declarations [38]. Most instructive is the shape of the remedy: OpenODD deliberately separates the attribute taxonomy (delegated to ISO 34503 and PAS 1883) from the declaration language, requires every situation to be decidable binarily as in-domain or out-of-domain, and stipulates that declarations remain valid for the vehicle’s full life cycle. In this section’s language, the industry is pulling merged layers apart: vocabulary separated from declarations, satisfaction semantics for membership, stability for certification artifacts. The theorem says a merged carrier must betray one of its goals; the casualty report records the betrayals as they actually occurred in practice — and the ongoing standardization effort independently re-enacts the separation prescription the theorem writes. ODD is not an isolated proof but the test paradigm of this proposition: any scheme that compresses the middle two layers into one carrier will see the predicted pathologies break out at the merger — a second casualty report is welcome. Witness two: an independent convergence on the same skeleton. The second witness comes from a purely mathematical tradition. Yuan and Yao’s calculus of intelligence (COIN) [39] formalizes agentic workflows as typed free-monad task spaces in a Grothendieck topos — a construction with no reference to the artificial physical world or to engineering archives. Yet measured with this section’s ruler, its structure has all four layers: from the monad composition laws to the type discipline, the certified blueprints, and the leaf contracts, each floor corresponds to one of this section’s goal pairs (the full inventory is in §8, R.6; the boundary of the applicability domain is in Appendix A.4). An independent mathematical tradition, without any reference to engineering archives, converges on the same four-layer skeleton. The two witnesses complement each other exactly: ODD testifies from the negative — merge and the pathology breaks out; COIN testifies from the positive — independence yields isomorphism. One symmetry must be acknowledged: the witnesses are independent of this paper; the translation of their testimony into this paper’s vocabulary is ours, and the selection was not blind — both were chosen because they map. The mapping is checkable, and its defeat criteria are stated where the translation is performed in full (§8, R.6). Witness three: layering convergence in the engineering lineage. Layering is not any one school’s preference either — beyond this paper, at least two independent sources in the engineering lineage point the same direction, though their layers are cut differently. Rasmussen’s abstraction–decomposition space has five levels, along the two axes of means–ends and whole–part, characterizing operators’ cognition [32]; Palantir’s Ontology [40] divides organizational modeling into semantic elements (objects, properties, links) and kinetic elements (actions, functions, dynamic security), serving data integration and operational closed loops [40]. What these practices share is layering; where they differ is the layering axes; and this section’s proposition explains exactly why the divergence: the division of layers is not arbitrary — it is uniquely determined by the construction goals; different goals, different axes. This is testimony one grade weaker than ODD and COIN: it does not prove the lower bound is four; it proves that the mechanism “layering is forced by goals” recurs in independent engineering practice. The propositions of this section constrain only the skeleton’s shape; the skeleton’s engineering existence, and four demonstration cases — from the unrepairable Mars rover to the everyday chiller plant — are in Appendix B and Appendix C. 6. Falsifiable Predictions: Putting the Theory on the Table A story that only tells the past is history, not theory. If this paper’s framework is to deserve the name “research programme,” it must say things whose failure would damage the framework. The five predictions below are ordered from broadest to strictest; each comes with its theoretical basis and explicit falsification conditions. P1 (breakthrough order): AI’s capability breakthroughs will continue to advance along the constraint axis — after the symbolic world, the next scaled breakthrough occurs in the artificial physical world, not the open phenomenal world. Basis: the historical correlation of the constraint axis with the breakthrough order (§2.2, §3.2), and the dual-channel data interface of the artificial physical world (§3.2). Falsification: if the general world-model route (trained mainly on open phenomenal video) achieves scalably replicable success in physical operation before any “archive-driven” route — e.g., systems without semantic grounding reaching industrial-grade reliability in unstructured open environments — the explanatory power of the constraint axis suffers a heavy blow. P2 (the prior’s map of victory and defeat): in data-scarce design domains with readable archives, the prior-framework route will systematically beat the pure posterior route; in archiveless domains, prior extraction must lose. Basis: the legitimacy criterion and the readability gradient (§4.2, §4.3). Falsification: systematic counterexamples in either direction — a pure posterior method matching the cold-start performance of a prior framework in some archive-complete, data-scarce design domain (such as facility operations); or any hand-written content prior reviving in an archiveless domain. P3 (concept-layer invariance): the concept layer of a qualified prior framework is invariant across industries; the marginal modeling cost of onboarding a new industry decreases. This prediction is falsifiable, and testing it requires neither trusting us nor believing the theory — only consulting the public engineering record. The concept layer’s claim is: every legitimate purpose in the artificial physical world decomposes into the constitutive dimensions of artifactuality — transfer of energy, matter, information, or mechanical work, together with spatial relations; and every failure is an instance or combination of a closed failure-mode typology. Therefore any industry whose coherent engineering practice exhibits a purpose that refuses such decomposition, or a failure mode that refuses such instantiation, falsifies this prediction. If a falsifying observation exists, it lives entirely in the public record — engineering textbooks, design codes, standards, and failure archives — and any reader can go look. What is staked is not the finality of our current enumeration (Appendix B): it is the current best candidate, empirical and revisable, and spontaneous refinement within its vocabulary space neither falsifies nor corroborates this prediction. What is staked is whether the space of purposes and failure modes is closable at all — §5.4 gives the a priori argument (Proposition 5.1), and a single counterexample suffices to destroy it. P4 (the VLA ceiling): vision–language–action models (VLA) without semantic grounding [41; 42] hit a ceiling in institutional operations scenarios — demonstrative progress continues; institutional adoption stays absent. Basis: VLA faces the open phenomenal world; its priors are black-box and non-normative (§2.3, §7 O1 and its empirical note); institutional scenarios demand auditable, endorsable intelligence with reportable failures (§4.4, third leg). Falsification: VLA-class systems achieving endorsed deployment at scale in institutional settings such as facility operations or industrial inspection without connecting to any constitutive semantic layer. The threshold is stated, lest the condition be read as unfalsifiable: prompt strings, tool specifications, and orchestration rules are already priors by R.5’s criterion (§8), but they do not count as connection for this prediction; what counts is a domain concept layer with instance binding — types that denote, clauses anchored to assets. A VLA endorsed at scale with only such scaffolding falsifies the prediction. The rules of the game must be written out here, because this prediction has a position easily misread as a loss: if VLA gains institutional adoption by connecting to a prior cognitive framework, that is not the prediction’s failure but one of its modes of fulfillment — the wager is not whether VLA is used, but whether it can obtain endorsement alone. If a constitutive semantic layer stands behind the networked VLA, what stands there is still our theory. P5 (the irreplaceability of the ontology layer): progress in large language models will not replace ontology-based artificial intelligence that retains a prior cognitive framework — each capability jump of LLMs raises, not lowers, the institutional demand for a constitutive semantic layer. Basis: LLMs have types without instances, knowledge without a stance (§2.3, §7 O1); the semantic anchorage is the LLM’s complement, not its rival — the framework gives what the LLM reads a constitutive, auditable, instance-bound place to land (§2.3 position statement). The leading indicator of this prediction is designated not by market size but computed from the criterion’s gradient: practices that build ontologies for human organizations (e.g., Palantir) occupy the criterion’s middle band (§4.2 degrees of constitution, §4.3), where the prior’s standing holds only by half — if replacement happens, it must happen first at that most exposed position; if that position holds, the strongly constituted artificial-physical side needs no further argument. Falsification: generational progress of LLMs systematically bypasses the ontology layer — enterprises obtain endorsed decisions from general-purpose LLM agents directly in institutional scenarios such as operations, and the corresponding business of ontology-based companies shrinks accordingly. This prediction and P4 are two sides of one wager: P4 bets the posterior route cannot enter the institutional door alone; P5 bets that when it enters, it must bring us along. Appendix note: the concept-layer boundary of the organizational domain (a structural corollary, not a wager). Prior frameworks for human organizations exhibit a triple structure: (a) at the mission layer (the divergent zone), no cross-customer shared concept layer will appear — per-customer manual distillation (as in forward-deployed engineering) is that layer’s permanent cost; (b) at the convergent layer (finance, compliance, administration), a shared concept layer long ago exists and will continue to be reused (the ERP paradigm); (c) the framework’s value density rises with the codification of the organization’s mission — defense and intelligence highest, large industrial organizations next, ordinary commercial organizations lowest. Basis: organizations are the middle band where both conjuncts of the criterion are satisfied only by degree (§4.2 degrees of constitution, §4.3 middle band); competition, as the universal “physics” of commercial organizations, has anti-convergence as its first law — it produces shared adjectives (moats, niches) and forbids shared nouns (missions). Closing: Lakatos’s wager. The five predictions share one structure: each one’s failure corresponds to damage to a specific component of the framework — P1 damages the constraint axis, P2 the criterion, P3 the layering lower-bound proposition, P4 the normativity argument, P5 the complementarity structure of semantic anchorage and LLMs. We do not ask the reader to accept this programme because its arguments are elegant; we ask the reader to write down these five items and watch which way the world walks. By Lakatos’s distinction [43]: a research programme that busies itself patching the core when predictions fail is degenerating; one whose predictions keep cashing out and raising new questions is progressive. This paper chooses to put its programme on the table — if wrong, the framework is wrong. 7. Objections and Replies A theory paper’s honesty is measured by how it treats its own counterexamples. This section first does something rarely done: it records, on the books, one occasion on which the present framework overturned itself during its formation; then it answers six strongest objections one by one. Each objection appears first in its strongest version — weakened strawmen are not entertained in this section. Self-criticism: why not take “temporal priority of design” as the criterion. This theory indeed set out with the temporal criterion — where there is “design first, birth after,” the prior is legitimate. It was killed by a simple observation: genes also precede proteins. Temporal priority cannot distinguish drawings from genes, and so the criterion was upgraded to “intentional constitution ∧ readable archive” (§4.2). We record this self-overturning as it happened, because it is exactly a demonstration of how a criterion works: a good criterion is not announced; it is pruned by counterexamples. O1: Large language models have already read everything humanity has written — manuals, specifications, the text of drawings are all in the corpus. Isn’t the prior framework you discuss exactly what LLMs already have inside? This is the most dangerous objection the paper faces, and it deserves a four-part dismantling. First, the mode of digestion flattens normativity. An LLM reads a specification with the same machinery it reads a novel — the statistics of the next symbol. The “shall” in the constitutive document “supply-water temperature shall be 7∘C” is flattened by statistical digestion into “people usually write this”; the model can recite the norm, but has no mechanism by which it could be constrained by it. And normativity — purpose, function, fault — is exactly the only valuable ingredient in institutional scenarios. Second, emergent representations are not promulgated concepts. Mechanistic interpretability research shows that LLMs do contain decodable feature structure — probes and sparse autoencoders can read out entities, attributes, even crude world models. We do not deny it, but we disarm it first: such research reads regularities after the fact from an already trained machine, isomorphic to grammarians reading grammar from utterances and physicists reading laws from phenomena — interpretability research is the LLM’s physics, not its engineering. Concede one more step: interventional experiments (activation patching, steering vectors) prove these features are genuinely causal parts, more than after-the-fact storytelling. But even granting the features are all real, they remain: emergent, not promulgated — the next training run drifts them, and no one endorses their invariance; living in the interpreter’s vocabulary — the model itself cannot denote them; carrying no norm — a steering vector can make a model talk more about Paris; no feature can stipulate what it “shall” talk about, or what counts as a fault of deviation. The prior framework’s concepts are the opposite: they are undertakings that have been laid down — prior to data, constraining interpretation, lapsing on drift. What the other side found are the machine’s laws; what we erect are the world’s drawings — however true the laws, they cannot exercise the office of drawings. Third, types without instances. The LLM’s knowledge is an average of types — it knows what chiller manuals generally say, not whether this chiller at this moment deviates from its own specification. Symbol-to-symbol proximity is not symbol-to-world anchoring [7]. Fourth, not auditable. Institutional operations require state reports with determinate semantics (each level’s meaning fixed in the framework), not plausible prose. The difference between a gray box with a skeleton and a black box without one is, where endorsement is required, the difference between usable and unusable. In sum: the LLM has knowledge without a stance toward the world it has read. It has an honored place in our framework — the universal tool for reading archives — but it is a reader, not a layer; its place is below the four layers, not in the concept layer. One final note of industrial self-evidence. The engineering evolution of the large-language-model industry itself is a living counter-proof of this objection: every agent system that walks toward practicality erects human-promulgated scaffolding outside the model — system prompts, tool specifications, orchestration rules, memory files. This scaffolding is written to be read and constrains how the model shall act — by the criterion of Section 4, scaffolding is a prior. The full unfolding and lineage (direct descendants of electronic institutions) is in Section 8, R.5; here we register only this ironic fact: the pure posterior route’s own engineering practice is an existence proof of the necessity of priors. Empirical note. An internal evaluation by the first author’s team provides direct evidence for “types without instances”: 32 physics-consistency indicators designed for building-environment simulation tasks — covering temperature convergence, temperature-difference stability, start–stop sequencing, load response, setpoint–flow feedback, and other categories — were used to test five frontier general-purpose large models, with widespread failures; a system built on this paper’s architecture (prior semantic layer + posterior learning) passed essentially the entire indicator set. We register one structural fact: not one of these indicators is a subset of the laws of physics; every one is a subset of design intent — “the supply–return temperature difference shall remain stable” is not a thermodynamic corollary but a promulgation of operating specifications. A general model that has read everything humanity has written still cannot pass this examination, because the answers are not in the corpus; they are in the norms. O2: Three peripheral objections, settled together — the evolved object, the repurposed artifact, “isn’t this just ontology engineering.” First, AlphaFold succeeded on evolved objects — doesn’t that show priors can do without “intentional constitution”? AlphaFold is instructive for the criterion, but not for the reason “generation one used hand-crafted rules and failed, generation two switched to evolutionary data” (§4.3): both generations ingested the evolutionary record (multiple sequence alignments) as data — proving that conjunct (i) can be satisfied by happenstance, not that (i) is unnecessary. Note the price paid for (i)’s absence: there is no “intent” in the evolutionary record, and so AlphaFold can rank structures but can never say what a protein is for or what counts as its failure. Structure is accessible; normativity is not — and that is all a world not satisfying (i) can offer. Second, artifacts get repurposed — paperclips pick locks, screwdrivers pry lids, old factories become cafés; purposes drift, and design intent locks nothing down. Purpose drift is real, but it is no threat to the criterion, because drift’s mode of existence in the institutionalized world is not “silent deviation” but a new design event: an old factory becoming a café comes with renovation drawings, change-of-use approvals, and fire-safety re-review — intent is re-promulgated, the archive re-issued, and normativity changes hands wholesale with the new intent. Non-institutionalized repurposing (the lock-picking paperclip) exists, but it does not belong to the claimed domain of “high utility value, institutionally maintained” (§4.3) — the framework never promised to police every paperclip. Third, isn’t this just ontology engineering? Ontology engineering supplies the craft; this paper supplies the license. Ontology engineering since Gruber answers “how to build an ontology properly” but not “over which world is building an ontology legitimate, and why” — engineering practice presupposed legitimacy; this paper gives its criterion (§4.2), its cost map (§4.3), its sufficient conditions (§4.4), and the shape a qualified ontology must satisfy (§5.4, the layering lower-bound proposition). If this is ontology engineering, then it is ontology engineering that has finally explained its own presuppositions. O3: Drawings govern construction, but not use and aging. Real operations knowledge is tacit, living in the veteran masters [27] — buildings “learn” by themselves; deviation from design is the norm. The scope this objection marks out we partly accept, partly dissolve. The accepted part: the criterion claims only the range “recorded in design” — the framework honestly does not cover purely tacit knowledge (the §4.3 guardrail). The dissolved part has four pieces. First, failures in use and aging are not beyond rules — the failure evolution of single devices and spaces can have its source categories systematically enumerated by the excitation structure “man, machine, material, method, environment + system inputs”; what is tacit is the concrete craft, not the semantics of failure. Second, “unmeasurable phenomena” are in the overwhelming majority superpositions of known failure modes, not new concepts (§5.4’s disenchantment corollary). Third, institutionalized operations is precisely the historical process of progressively codifying tacit knowledge — the framework is that process’s carrier, not its enemy. Fourth, the harder form of this objection — archives that are not merely silent but systematically unfaithful: as-built diverging from as-designed, expired drawings, models detached from the site — is the case the criterion is built to make loud. Under the direction of fit, a divergence between archive and artifact is itself a violation event, loud and locatable, not a silent model error; and §4.4’s detection–quarantine–ratification protocol is the standing machinery that brings a repaired clause back across the waterline. An unfaithful archive is therefore a cost with a known price — archive completion is itself an engineering activity (§4.4) — not a counterexample to the criterion. O4: In complex, tightly coupled systems, unanticipated interactions among multiple faults can trigger system-level accidents [44] — do such accidents exceed the capability of any component-level prior framework? Before answering, we must ask the other side for a premise: faced with a system-level accident, what is intelligence’s task? Three tasks must be accounted separately. First, prior prediction. A system-level accident genuinely formed by the symptomless convergence of multiple independent factors indeed cannot be predicted in advance — but this is not our framework’s particular incapacity; it is any intelligence’s structural incapacity: such events belong with earthquakes and wars, force majeure, moved out of the domain of claim by §4.3’s demarcation corollary. To use them as counterexamples is to blame an insurance contract for not covering war. Second, detection and handling in the event. The overwhelming majority of so-called system-level “emergence” is not emergence at all but the stacking and cross-system spread of small errors: Chernobyl was not an indecomposable holistic catastrophe but a chain of individually observable, identifiable, explicable small errors — the reactor’s positive-void-coefficient design defect, the mismatch between operating procedures and physical characteristics, the safety test conducted in violation — each link a failure specifiable and classifiable in the technical semantics of its time — each falling within the semantics of known failure types. And the propagation of failure between systems can itself be modeled: systems define standard semantic interfaces between one another, transmitting graded states rather than physical parameters (Appendix B) — every link of the propagation chain is then a monitorable interception point. An intelligence that can recognize each small error on the spot as a “symptom” intercepts exactly the final accident. Third, post-hoc attribution and explanation. This is the framework’s home ground: ascending the collapse chain, every judgment gives reasons, every failure reports a location (§5.4) — the attribution report is a byproduct of the structure, not extra engineering. Add one symmetry: the posterior route has no advantage here — a black-box model facing an unseen common-cause failure is equally extrapolating, and its failure is silent (§4.4), while the framework’s failure is an alarm. Finally, put the disagreement on the table: P3 commits the concept layer to cross-domain invariance — if system-level phenomena truly require every industry to rewrite its concept layer, the facts will deliver the verdict on this objection for us (§6). O5: Physical laws work as priors in the basic physical world — PINNs carve Navier–Stokes into the loss function and succeed. So descriptive rules are equally legitimate priors, and conjunct (i), “intentional constitution,” is simply unnecessary. This objection’s edge comes from a category mistake: it takes “explicitly expressed posterior” for “prior framework.” Physical laws are not norms the world was promulgated before its birth but humanity’s statistical summary of centuries of observation of the basic physical world — their epistemic identity is the same as grammar’s: descriptive rules, posterior to phenomena (§4.1). What PINNs do is pre-position extremely compressed posterior knowledge as a training constraint — engineering-wise clever, but it supplies no constitutive norms: no “fault” semantics, only “prediction error”; no inheritance of promulgative force — when observations disagree, it is the law that gets revised, not the world (§4.4). Its legitimacy is therefore retrospective and unrelated to our criterion; its scope is strictly limited to the small set of problems “approximable as closed physical systems” (§4.3). A physical law is the posterior written out explicitly, not a prior; it describes how the world runs, not how the world shall be maintained. The complementary division of labor — physics engines supply constraint boundaries, the constitutive semantic layer supplies normative coordinates — and the surgical comparison with this nearest neighbor, are in Section 8, R.7. O6: Where is the learning theorem? This paper speaks of legitimacy, promulgation, and direction of fit, yet proves nothing a learning theorist can use — no sample-complexity bound, no generalization guarantee, no PAC-style separation from Bayesian or learned priors. What does any of this change for machine learning? The fact is conceded at once; the charge is contested. The charge presupposes that a “prior” is a well-defined object whose only open question is performance. The entire burden of Section 4 is that this presupposition fails in most domains: a prior framework is not merely a constraint on hypotheses but a claim backed by an authority source, and whether that source exists is logically prior to how well the claim generalizes. Legitimacy is therefore not a rival answer to the learnability question; it determines where the learnability question has an object at all. A bound proved for a prior that was never legitimate is a bound proved for a stipulation. What the criterion changes for learning is a classification, not a bound — and where the classification comes out positive, the central learning problem changes species. In domains satisfying conjuncts (i) and (i), the hypothesis space is not to be induced from samples but read off an archive; what remains is conformance checking against promulgated concepts — a strictly weaker task than induction, because the concept is given rather than inferred, and a confirmed deviation is a fault in the world rather than a revision of the model (§4.4). This, and not any accident of physics, is why the artificial physical world is the cheapest domain for machine intelligence: its concepts arrive pre-promulgated, and “learning” there means reading, instantiating, and checking. The demarcation section promised to add the learning-theory chapter to the sciences of the artificial (§3.4); that section’s first result is now statable: in the designed world, learning theory’s hardest question — induction — is replaced by its easiest — verification. Nor can any learning theorem supply the missing ingredient where the criterion fails. Induction mines regularities; it cannot mint the normative: no instance carries its own “shall,” and no sample can stand as a counterexample to a norm — the AlphaFold lesson of O2 is exactly that structure without normativity is all an unintended world offers. The computable deliverable of this paper is therefore a boundary, not a bound. Appendix A.4 states it as a decidable precondition on certificate-anchored composition calculi: before asking how well a prior-based method performs in a domain, ask whether the domain can ground the method’s antecedents at all; where it cannot, soundness holds at most vacuously, and certificate chains degrade into stipulations. The wager just made is weaker than a theorem and stronger than a theorem: weaker, because no bound is proved; stronger, because the paper claims to delimit the territory inside which questions about prior-based learning are well-posed at all — and has staked that claim, in Section 6, on the public record. 8. Theoretical Positioning and Related Work This section settles accounts with three kinds of traditions: those we draw from, those converging with us, and those from which we must demarcate. The structure is deliberate: one section of demarcation and docking (R.1), four independent strands of convergence evidence (R.2–R.5, including the two demarcation sections R.2’ and R.3’), one parallel mathematical construction inventoried through this framework’s lens (R.6), one control group isolating the key variable (R.7), and one closing section (R.8). We did not invent the convergence recorded here; this paper is its theory. R.1 The KR Tradition and Description Logics: Shared Conceptualization, or Constitutive Promulgation? Knowledge representation is the fifty-year institutional home of explicit semantic structure in AI. But its founding definition carries a neutrality that later became a blind spot. Gruber’s definition of ontology — “a specification of a conceptualization” [14] — is deliberately silent about the epistemic status of that specification: it may be law promulgated to the world, or a summary extracted from the world; the definition prices both the same. Section 4’s criterion is exactly the question this neutrality suppressed: did the object domain itself leave a readable constitutive archive, prior to its instances? The tradition’s most self-conscious branch has recognized that concepts need meta-theoretic discipline. OntoClean’s meta-properties — rigidity, identity, unity [45] — constrain what a concept may legitimately mean, and Guarino’s formal ontology [15] anchors such constraints in philosophical analysis. But the source of the constraint matters. OntoClean’s discipline is descriptive: it codifies how philosophers and careful modelers ought to analyze concepts; its authority rests on the persuasiveness of argument. Our concept layer is promulgated: its constraints are promulgated by the design archives of the artificial physical world; its authority rests on the promulgative force of norms. The KL-ONE tradition [46] pushed the structural-description view of concepts to its extreme — a concept is a node in a subsumption lattice — yet never touched: by what right may this lattice exist prior to experience? CYC [1] is the tradition’s most ambitious wager on content priors, and its failure lies not in the representation language but in the choice of domain: common sense is an archiveless domain (§4.3). By contrast, McCarthy’s circumscription programme [47] addresses a different problem — how to draw defeasible conclusions from incomplete common-sense axiom systems; our diagnosis therefore targets not circumscription itself but the broader ambition of pre-specifying domain-general common sense in the absence of a readable constitutive archive. The demarcation in one sentence: the KR tradition built magnificent normative machines but never asked whether the domains it chose left readable constitutive archives; the criterion is that question, posed and answered. Description logics deserve a table of their own. They provide the closest formal counterparts to our four-layer skeleton, precise enough to tabulate: This paper DL tradition syntax layer signature (vocabulary) concept layer TBox (terminological axioms) knowledge layer rule base (SWRL-style rule extensions) instance layer ABox, extended with streaming data This table is no analogy: our engineering implementation is realized in OWL/RDF, and the vertical edges of §5.4 are literally inheritance and instantiation relations in the implementation. But the two traditions arrive at layering from opposite directions, and that difference is the contribution. In the DL literature, the TBox/ABox separation is tradition rather than theorem: inherited and refined as representational practice, with expressiveness choices closely tied to reasoning complexity and decidability [31; 48; 49]. Our layering lower-bound proposition (§5.4) gives, to our knowledge, the first argument that at least four separations are enforced by the construction goals themselves — stability versus openness on rate of change, abstraction versus concreteness on degree of abstraction, four mutually incompatible goal pairs, hence at least four layers. The driving variable is not computability but maintainability: when everything is growing, what must stand still. The two bodies of results complement, not compete. DL tells you how expressiveness choices change reasoning complexity; we tell you why any framework that “both grows and remains one system” needs at least four separations. KR engineers may read our proposition as the first engineering justification of their long practice; we read their decidability results as the mathematical insurance policy for the layers our proposition enforces. R.2 Normative Systems: The Oldest Formal Home of Promulgation The normative-systems tradition, looked at closely, is our closest theoretical ally — the community that formalized promulgation decades before the current debates. Deontic logic asked from the start: what is the logic of “ought” [50]? Normative multi-agent systems turned the question into engineering: norms have life cycles — creation, promulgation, compliance, violation, sanction — and not one link is statistical. Obligations are promulgated by legislators, not averaged from behavior; violation is followed by sanction, not by a silent update of a distribution. Electronic institutions [51] made it concrete: formalized, computable norms for agent societies, with roles, interaction protocols, and normative rules all specified before agents participate. BDI agents [52] and practical agent languages such as 2APL [53] internalize explicit motivational and deliberative structure: goals, intentions, and plans are stipulated as architectural components, not implicit in observable behavior. A terminology debt must be paid face to face. Boella and van der Torre distinguish regulative from constitutive norms in normative MAS [54] — same word as ours, same Searlean lineage as ours. The difference lies in what is constituted. Their constitutive norms constitute institutional facts within agent societies (“this utterance counts as a promise”); ours constitute the normal states of the artificial physical world (“this supply-water temperature counts as within specification”). They legislate for agent societies; we legislate for the world of artifacts. Isomorphic structure, different objects. The exchange between the two traditions runs both ways. Normative MAS presupposes promulgative authority — that a legislator can legislate prior to behavior — and our criterion answers when this presupposition holds: in domains intentionally constituted and documented with readable archives. In return, their formal machinery (deontic operators, norm life cycles, enforcement models) provides a ready-made formal language for our constitutive norms. The scaffolding of contemporary LLM multi-agent systems — roles, protocols, orchestration rules — is the direct descendant of electronic institutions (R.5); the LLM industry is rediscovering this tradition in the wild, and both communities deserve to know it. R.2’ Bayesian Priors: A Terminological Demarcation The Bayesian tradition places, before data arrives, a probability distribution over parameters or hypotheses — a subjective prior derived from the modeler’s beliefs, updated by Bayes’s rule, and by design washed out as evidence accumulates. Its similarity to this framework is terminological only, and the difference is exact. A Bayesian prior is a quantitative belief: its content is a distribution, its source the expert’s head, its direction of fit from world to belief — when data conflicts with the prior, it is the prior that moves. A constitutive prior is a promulgated norm: its content is a framework of types, rules, and clauses; its source the design archive; its direction of fit from norm to world — when an instance deviates from the norm, the instance is at fault, and the framework’s duty is precisely to say so. The two also fail differently: a misspecified Bayesian prior is diluted by the likelihood, degrading gracefully and silently; a violated constitutive norm goes off with a bang at a named clause. Formally, the two do not live in the same mathematical space: one lives in probability-measure space, the other in the model-theoretic satisfaction relation (⊧ ) — an instance either satisfies the framework or it does not, and no likelihood ratio can answer that question. This is the formal echo of the is–ought gap: a distribution over “what happens” cannot express “what shall happen.” The subjective prior is no adversary here — the instance layer may legitimately hold such priors for quantities the archive leaves unspecified. But no amount of Bayesian updating can turn a distribution into a norm: the posterior remains belief, and belief does not bind. R.3 Industrial Ontologies: Engineering Forced into the Semantic Layer by Pain The engineering world independently supplies a striking existence proof: layered, promulgated semantics is not a philosophical taste but a practical necessity — and its producers have read no philosophy. Three signals deserve registration. Palantir’s Ontology, a representative deployed industrial ontology, organizes organizational modeling into semantic elements — objects, properties, links — and kinetic elements — actions, functions, and dynamic security [40]. Its layering divergences and forward-deployed cost structure we analyzed in §5.5 and §6; here we register only the fact: one of the most watched industrial data platforms of the decade has, at its core, a promulgated semantic layer with an operational closed loop. Google’s knowledge graph teaches the same lesson at web scale — “things, not strings” [55] — but by this paper’s framework it is a descriptive product, belonging to §4.3’s “happenstance-readable” class, not the constitutive one. The third signal is the most instructive, because it comes from the grassroots of our own industry. The building-automation industry — an industry with no AI-theoretical ambition whatsoever — was forced by the pain of integrating mutually incompatible building-control data into inventing its own semantic layer: Brick Schema [56] defines a unified ontology for building entities and their relations; Project Haystack [57] maintains a community-governed ontology of equipment and data points with controlled tag vocabularies — both promulgated prior to all analytics. No one in this industry was persuaded by theory; they were cornered by the economics of integration. When a domain’s data cannot interoperate without shared semantics, the domain invents shared semantics — promulgated, versioned, committee-maintained semantics. The criterion explains why this invention is compulsory, not elective; the layering lower-bound proposition predicts its layered shape. The same verdict covers the digital-twin wave: a twin without a promulgated semantic layer is a dashboard, not a model of the domain. R.3’ Safety Boundaries: ODD and SOTIF Automotive safety regulation independently converged on a promulgated normative boundary: the Operational Design Domain (ODD) [35; 36], with its edge behaviors governed by the Safety of the Intended Functionality framework (SOTIF) [58]. An operating instance either satisfies the declared domain or violates it; leaving the domain triggers a prescribed minimal-risk response, not a model update — the direction of fit is normative by law. We read it as a specimen, not an adversary: a constitutive structure rebuilt by an industry with no reference to this framework. What this practice lacks, and this paper supplies, is the logical foundation: satisfaction semantics for domain membership (⊧ ) (R.2’), decidability of boundary checks over closed attribute vocabularies (Observation A.6), and clause-level accountability (Section 5). SOTIF’s partition of scenarios into known-safe, known-unsafe, and unknown-unsafe is exactly the applied form of the novelty–violation distinction (Remark A.5’): the former shall be learned; the latter must be stopped. R.4 Neuro-Symbolic AI: Our Closest Fellow Traveler in the Learning Community, and the Foundation It Lacks For over two decades, neuro-symbolic AI has persisted against the wind: statistical learning alone is insufficient for robust and interpretable intelligence; neural learning must be combined with symbolic knowledge representation and logical reasoning [59]. Every time an unstructured model capsizes on a systematic, auditable task is a belated endorsement of this insistence. We regard this community as our closest fellow traveler inside learning science, and “close” is literal: their engineering results constitute indirect evidence for our criterion. Every gain from symbolic injection — the sample efficiency of the Neuro-Symbolic Concept Learner [60], the probabilistic-logic integration of DeepProbLog [61], the differentiable logical constraints of Logic Tensor Networks [62] — is a case of prior structure doing work. Our criterion explains why injected structure works: it works to the exact degree that it plays promulgation rather than description. A division of labor follows. Neuro-symbolic research answers the architecture question — how symbols and networks couple: injection, parallelism, serial distillation; there may be many answers, and we place no bet in that race. This paper answers the prior question — where symbols come from, and by what right they constrain networks. The coupling style is a free design choice; promulgative authority has exactly one source: intentional constitution with a readable archive (§4). What the neuro-symbolic tradition lacks is a defense of its symbolic layer’s epistemic status — why these symbols may precede the data, and who vouches for their normativity. Our criterion and lower-bound proposition are that defense. To our closest fellow traveler we offer not criticism but a title deed: every neuro-symbolic architecture that anchors its symbols in constitutive archives is, by our criterion, legitimate engineering of prior intelligence. R.5 Agent Scaffolding: The Return of Promulgation in the Wild The LLM industry’s own engineering evolution is the most contemporary of our four strands of testimony. Its trajectory, step by step, is a history of promulgation returning through the back door: prompt engineering became system prompts; system prompts grew tool specifications [63; 64]; tool specifications were joined by persistent memory stores [65]; multi-agent systems added orchestration rules — roles, protocols, division of labor [66]. Every step: posterior capability proved insufficient, and a promulgated artifact was patched on. These artifacts stand up to a close reading with our criterion. They are not descriptions distilled from model behavior but norms written by humans to constrain model behavior; they are written to be read — indeed they occupy the most extreme cell of our archive gradient (§4.3): the only documents in existence whose first reader is a machine. Their direction of fit is world-fits-document: the whole value of a system prompt lies in its promulgating how the model shall act, not in describing how it usually acts. By the criterion, scaffolding is a prior: intentionally constituted, written to be read. The industry’s own practice thus becomes an existence proof of the necessity of priors — offered by the purest posterior tradition in AI’s history, not without irony. The keywords of this new engineering are not “prior” and “posterior,” but “promulgated” and “described.” R.6 A Certified Agent Calculus: An External Inventory A parallel mathematical tradition confirms the four-layer lower bound from an independent direction. Yuan and Yao’s calculus of intelligence (COIN) [39] formalizes agentic workflows as typed free-monad task spaces in a Grothendieck topos. Inventoried through the present framework’s lens, COIN has all four layers: the workflow signature and the monad composition laws perform the syntax layer’s office (stable, abstract); type declarations in the typed task space (e.g., UserID, AuthResult) constitute the concept layer (open, abstract) — they regulate what plans are admissible, not what the world happens to contain; the per-task generated, compositionally certified and finalized decomposition trees and compiled blueprints — together with their certificate chains — constitute the knowledge layer (stable, concrete); and real-world instances land only indirectly, through leaf contracts and external implementations, constituting the instance layer (open, concrete). One terminological point deserves precision: a bare Kleisli arrow is a morphism — a description of “how to compute from I to O with effects” — not a static classification; by itself it instantiates no layer. What performs the concept layer’s office is the type discipline constraining arrow admissibility; only when arrows are composed, certified, and finalized into blueprints does the stable × concrete layer emerge. An independent construction, without any engineering archive, converges on the same four-layer structure — this supports Section 5’s lower-bound proposition. The boundary is equally instructive. COIN’s instance layer docks onto external human-run systems — databases assumed to realize declared schemas, session services assumed to satisfy stated contracts, and, in its organizational-planning example (their §6.3), budgets and compliance requirements — systems that, in this framework’s language, are intentionally constituted and documented, the organizational ones with considerable grayness in both norms and documentation (§4.3). There, certificates can treat external systems as stipulated hypotheses, because the legitimacy of such declarations is usually a matter of convention. The artificial physical world offers no such shortcut: engineering archives carry liability, not convention, and the legitimacy of extracting priors from them must be answered head-on — which is exactly the question Section 4 exists to answer. It is a coincidence worth noting that COIN’s own organizational-planning example is acknowledged by its authors to be “less formal” than the theorem-proving and software cases — an independent observation of the same gray band. The criterion also predicts where the certification calculus falls silent. COIN’s calculus presupposes already-promulgated symbolic structure at every step: differentiation requires typed tasks whose types are drawn from existing schemas and contracts; integration requires certificates anchored to declared assumptions; and the framework’s own applicability gate admits tasks “only after [their] relevant context has been made explicit” (their §2.1). In worlds with neither promulgated norms nor readable archives — unstructured terrain, weather, geology — type declarations can only be stipulated, and certificate chains hang in the air. This is not a defect of COIN but a prediction of the legitimacy criterion, made precise as a boundary theorem in Appendix A.4. The prediction is falsifiable: attempts to deploy certified decomposition calculi in archiveless environments will see their certificate chains degrade, at the boundary, into unverifiable stipulations. For the record: this framework was developed independently of COIN; the inventory above was made after the lower-bound proposition was established, and is offered as an external test of it. R.7 Control Group: PINNs, and the Two Fates of Explicitness Physics-informed neural networks [22] are the one neighbor that must be separated with surgical precision, because they resemble us in every surface feature: explicit structure, distrust of pure black boxes, respect for the domain’s laws. That resemblance is exactly what makes them this section’s control group. Fix the variable “explicit” — both sides have it — and flip the variable “promulgated”: PINNs’ explicit content is descriptive (laws extracted from observations of the world’s behavior); ours is constitutive (norms promulgated about how the world shall be maintained). The consequences fork accordingly. Physical laws describe how the world runs; they do not define how the world shall be maintained — “the supply–return temperature difference shall remain stable” is not a subset of any physical law (§4.3, §7 O5). PINNs inherit the full costs of the archiveless world they compress: their norms have no failure semantics, their silence toward deviation is silent failure, and their legitimacy is retrospective. The relation is complementary, not competitive; §7 O5 has registered the division of labor: physics engines supply constraint boundaries (what is impossible), the constitutive semantic layer supplies normative coordinates (what counts as normal, what as fault). But for reviewers and readers, PINNs provide a scalpel for identifying us against our nearest neighbor: do not ask whether the structure is explicit; ask whether it is promulgated — deviation from it: a violation in the world, or a revision of the model. A companion demonstration of this distinction — a training process guided by a prior semantic network on building thermal data — is currently in preparation. R.8 Closing: Four Communities, One Control The evidence gathered in this section has only one shape. Four communities — engineering (R.3), learning science (R.4), normative-systems theory (R.2), and the LLM industry itself (R.5) — plus one parallel mathematical construction (R.6) — driven by different pains, working on different objects, sharing no agenda, have independently converged on the same structure: an explicit, promulgated, layered semantic layer. One control group (R.7) isolates the decisive variable: what matters is not explicitness but promulgation. We did not invent this convergence; this paper attempts to be its theory — to say what the four communities voted for with their feet, why the vote was right, and where its boundary lies. With the criterion as instrument, the routes to the physical world measure as follows: Route Data source Direction of fit Failure semantics Zero-shot operable deep learning instances descriptive silent drift no world models video, interaction descriptive (prediction) silent no PINNs observations + laws descriptive (explicit) prediction error, not violation no knowledge graphs extracted concepts descriptive none partial rule era (Cyc, expert systems) hand-written rules promulgative — wrong domain loud but brittle “yes” where there are no rules constitutive priors (this paper) design archives promulgative loud violation yes Table 2: The routes to the physical world, measured with the criterion as instrument. 9. Conclusion: The Artificial Physical World — Low-Hanging Fruit on the Road to AGI Return to the deadlock we started with: no intelligence without data, no data without intelligence. All the work of this paper can be closed into one answer to that deadlock — the deadlock is universal, but not uniform: civilization left a crack in it, in its most valuable artifacts. That crack is called design. The derivation of the four worlds tells us the cognitive world is layered for the learner; the criterion tells us that among the four floors only one permits “framework first, instances after”: the artificial physical world, intentionally constituted and documented with readable archives; the layering lower-bound proposition tells us such a framework needs at least a four-layer skeleton — the syntax layer promulgates the norms, the concept layer holds stability, the knowledge layer stores experience, the instance layer lets posterior data land, and the collapse semantics lets the complicated reduce upward. The flywheel thereby replaces the deadlock: a prior framework gives intelligence that runs from day one; running produces data; data flows back at the instance layer; intelligence grows in use. This does not mean the other three worlds can be bypassed. The phenomenal world, the basic physical world, the artificial symbolic world — each is an inexcisable floor of a complete intelligence; although this paper has not proposed a path to general intelligence, if AGI is to stand, no floor can be missing. But the floors play different roles: only in the designed world can intelligence’s understanding of the world be reviewed layer by layer, interrogated clause by clause, endorsed level by level. The intelligence this paper advocates does not pretend to be transparent down to every gram of flesh — its flesh is posterior machine learning, and no one can interrogate every parameter; it is a gray box. But a gray box has a skeleton: the skeleton is prior promulgation, deciding where the flesh grows, how it grows, and where it alarms when it grows wrong. Models of the phenomenal and basic-physical floors forever carry the risk of silent out-of-distribution failure; the giant of the symbolic floor has knowledge without a stance toward the world it has read; only the cognitive framework of the artificial physical world, because it inherits promulgation rather than extrapolation, can trace every judgment up the collapse chain into reasons and locate every failure at a specific layer. In an age when more and more decisions are handed to machines, intelligence that can be interrogated is not a luxury but infrastructure — and interrogability comes not from transparency but from the skeleton. We call it interrogable: every one of its judgments can be pursued along the collapse chain to a promulgated normative clause, and where it cannot answer, it alarms — not interpretable (an explanation can be an after-the-fact fabrication), and not merely auditable (an audit inspects records; an interrogation demands answers in court). A neighboring engineering tradition has pursued a cognate goal for two decades: safety cases and assurance cases [74], most recently extended to machine-learned components (AMLAS [75]), demand exactly what we demand — that every claim about a deployed system be pursued to explicit evidence. The tradition supplies the engineering method but has lacked an ontological ground: an assurance case argues that a system is safe; it does not say what kind of world makes such argumentation possible at all. This paper’s criterion is the missing ground — assurance cases are buildable exactly where worlds are intentionally constituted and readably archived; outside the criterion’s boundary, assurance arguments degrade into the unverifiable stipulations of Proposition A.4’s boundary theorem. Which world ships with its own source code? — only the designed one. This floor of civilization lays its drawings openly on the table; learning to read them is the most honest, and the most economical, door for machines entering the physical world. Two questions this paper deliberately leaves open are taken up in two companion papers: what form the intelligent body should take in the artificial physical world [69], and what kind of carrier can build and run this framework’s knowledge and instance layers [70]. The three papers share a single axiom — the artificial physical world is purpose-first. Appendix C organizes the framework’s demonstrations as a register of four cases in two tiers — the unrepairable Mars rover and the everyday chiller plant walking the full reasoning chain; a factory’s pure-water system and a city’s heat network carrying failure-mode instantiation statements — with the register’s selection rationale, disclosure status, and falsifiable commitments stated on the record there. Acknowledgments. The authors thank their colleagues Wenjia Gao and Jieshi Xiao at Persagy Science and Technology for discussions on the engineering practice underlying Section 5 and Appendix B. Declaration of generative AI and AI-assisted technologies in the writing process. During the preparation of this work the authors used Kimi (Moonshot AI) for editorial assistance, including language polishing, Chinese–English translation, structural-editing suggestions, and consistency checking of the manuscript. All concepts, arguments, analyses, and conclusions are the authors’ own. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication. Appendix A. Semi-Formal Statements of the Criterion and the Lower-Bound Propositions This appendix writes the paper’s three core arguments as semi-formal statements: the legitimacy criterion of §4.2, the promulgation constraint and direction of fit of §4.4, and the layering lower-bound proposition of §5.4. The purpose is to give reviewers a formal skeleton checkable item by item; the “formality” stops at the level of predicates and propositions and claims no axiomatization — the paper’s argumentative burden is philosophical analysis, not logical derivation. A.1 Basic Conventions Definition (A.1 — object domain). An object domain D is a triple ⟨ , N, A⟩ : • I: the set of instances — observable, identifiable objects and processes in the domain; • N: the set of norms — statements about how instances in I “shall be”; • A: the archive — records existing in symbolic form, readable by machines. Definition (A.2 — constitution). A norm set N constitutes an instance set I, written N ⊲ I, if and only if: the existence and identity of every instance in I depends on some set of norms in N being complied with — that is, “this instance being this instance rather than another” is stipulated by norms. Constitution satisfies priority: N is fixed prior to the production of the corresponding instances in I. Two semantic demarcations against misreading: ⊲ is not the necessity operator of modal logic — it says nothing about possible worlds; nor is it concept subsumption in description logics — what it relates is not concepts to concepts but norms to instances. Its semantics anchors in the act of promulgation: an intentional promulgation stipulating what instances shall be (§4.2), in the lineage that the normative-systems tradition formalized long before this paper (Section 8, R.2). Definition (A.3 — promulgation and description). For a norm n ∈ N and an instance i ∈ I: • if n ⊲ the relevant instances and an i violating n is judged a defect of the world (i is judged a fault, a violation, to be corrected), then n’s direction of fit to i is promulgated; • if an i violating n is judged a defect of the model (n is judged a misfit, to be revised), then n’s direction of fit to i is described. (The direction is the direction of fit in the Anscombe–Searle sense: promulgated = world-to-word; described = word-to-world.) Definition (A.3’ — violation resolution). Let d(i, n) be the event of instance i deviating from norm n. The direction of fit determines the resolution operator. If n is promulgated, d(i, n) resolves to Violation(i) — a defect of the world: locatable, reportable, correctable against the promulgated clause. If n is described, d(i, n) resolves to ModelDefect(n) — a defect of the model: absorbed by further learning into parameter revision. The two resolutions are not interchangeable labels but two different event types, with different downstream operators attached: the former issues an alarm with an address (§4.4); the latter silently updates a distribution. Definition (A.4 — readable generative archive). An archive A is a readable generative archive of D if and only if: (i) A records (part or all of) the promulgated content of N; (i) A exists prior to the corresponding instances; (i) A exists in machine-parseable symbolic form. A.2 The Legitimacy Criterion (Formal Skeleton of §4.2) Proposition (A.1 — the criterion). For an object domain D = ⟨ , N, A⟩ , extracting a prior cognitive framework over D is legitimate if and only if: • (C1 intentional constitution) N ⊲ I — the domain’s instances are constituted by prior norms; • (C2 readable archive) there exists a readable generative archive A of D. Argument (necessity). Suppose C1 fails: instances are not constituted by prior norms; then the content of any “prior framework” can only come from induction over existing instances, and its direction of fit is description — the “framework” is in fact a product of posterior learning, prior in name only (§4.3: physical laws relative to the phenomenal world are exactly this case). Suppose C2 fails: norms constitute instances but left no readable record; then the framework cannot be extracted, and legitimacy has no engineering ground (§3: the case of the basic physical and phenomenal worlds — constraints objectively exist but have no archive). ∎ Assumptions (premises for the sufficiency argument only). A1 (archive fidelity): the archive faithfully records the currently promulgated norms — maintained and versioned; a gap between archive and promulgation is itself a correctable event. A2 (semantic accessibility): the learner can correctly interpret the archive’s vocabulary — parseable and groundable. A3 (coverage): judged instances fall within the coverage recorded by the archive; beyond coverage, the verdict is “unknown,” never a silent error. Proposition (sufficiency, restated). If D satisfies C1 and C2, then under A1–A3, a framework F extracted from A satisfies S1, S2, S3: • (S1 promulgative constraint) F’s constraint on I has the direction of promulgation: an instance’s deviation from F is a defect of the world, not of the framework; • (S2 loud failure) by S1, deviation necessarily triggers a “violation” verdict rather than a silent misfit — failure is locatable and reportable (§4.4’s “alarm with an address”); • (S3 zero-shot operability) F is fixed prior to instance data, so a system loaded with F has judgment capability before the first instance arrives — the deadlock (§1) is untied. Corollary (A.1 — verdicts on the four worlds). Testing each domain with Proposition A.1 (the three-variable table of §3): the phenomenal world fails C1 (no prior constitutive norms); the basic physical world fails C1 (laws are descriptions, not promulgations); the artificial symbolic world satisfies C1 while C2 holds partially, but its norms do not bind physical instances (§3.4, the source-code partition); the artificial physical world satisfies C1 and C2 simultaneously — it alone is the object domain of legitimate prior frameworks. A.3 The Layering Lower-Bound Proposition (Formal Skeleton of §5.4) Definition (A.5 — construction goals). Building a prior framework pursues four goals (§5.2): • G1 stability: the framework’s core does not drift with the instance stream; • G2 openness: the framework can ingest unforeseen new instances and new domains; • G3 abstraction: compress instances into concepts — accumulable, iterable, reusable; • G4 concreteness: descend to anchor individual instances, supporting on-site judgment. Definition (A.6 — carriers and goal pairs). A carrier of a framework is a content region with a unified revision discipline and a unified denotation discipline (§5.3). A carrier v realizes a goal pair (Gx, Gy) if v’s structure satisfies both goals’ requirements. Proposition (A.2 — incompatibility). The four goals form two pairs of contradictions: • (P1) G1 and G2 are mutually exclusive in a single carrier: a carrier open to modification by all new instances cannot remain stable; locked for stability, it cannot be open. • (P2) G3 and G4 are mutually exclusive in a single carrier: a fully abstract carrier loses instance-level anchors; a fully concrete one loses compression and reuse. The six pairwise mergers of the four goal pairs, checked item by item — every merger triggers at least one discipline conflict: Merger Shared Violated (S×A, O×A) abstraction revision discipline (freezing vs ingestion) (S×C, O×C) concreteness revision discipline (S×A, S×C) stability denotation discipline (type vs instance) (O×A, O×C) openness denotation discipline (S×A, O×C) none doubly violated (O×A, S×C) none doubly violated Proposition (A.3 — layering lower bound). A prior framework satisfying the criterion of Proposition A.1 needs, on the two definitional dimensions of time (promulgation precedes and constrains the instance stream) and denotation (constraint reaches instances as type or as instance), at least four carriers, realizing the goal pairs respectively: • L1 syntax layer ⟵ (G1, G3): stable and abstract — promulgated vocabulary and connection rules, not drifting with instances; • L2 concept layer ⟵ (G2, G3): open and abstract — domain concepts augmentable along inheritance, abstraction unbroken by openness; • L3 knowledge layer ⟵ (G1, G4): stable and concrete — instantiated knowledge codified clause by clause, frozen thereafter; • L4 instance layer ⟵ (G2, G4): open and concrete — continuously flowing posterior data and individual instances. Argument (reductio). Suppose the framework uses only k < 4 layers. Then by the pigeonhole principle, at least one carrier must realize goals from two mutually exclusive goal pairs, or both ends of P1 / P2 simultaneously. By Proposition A.2, the two ends cannot be held by one carrier — that carrier either cedes stability to openness (the framework drifts, violating G1), or cedes concreteness to abstraction (losing anchors, violating G4), or the reverse. Hence k ≥ 4. ∎ Remark (A.1 — lower bound, not upper bound). Proposition A.3 proves no further compression (below four layers, some pair of goals must injure each other); it does not prove four layers sufficient — engineering implementations may subdivide within layers (Appendix B’s six object classes are the concept layer’s internal structure). If future practice requires a fifth or sixth layer, a separate paper will argue it; this paper’s claim stops at the lower bound. Remark (A.2 — uniqueness of inter-layer relations). There is only one kind of inter-layer relation: abstraction and instantiation (parent–child sets, taxonomy) — knowledge-layer instances instantiate concept-layer concepts, concept-layer concepts instantiate syntax-layer word classes, instance-layer objects instantiate knowledge-layer entries. Top-down it is instantiation layer by layer; bottom-up, collapse layer by layer; the reduction path is the explanation path (the formal core of §5.4’s interpretability promise). Remark (A.3 — on the discreteness objection). Reading the two goal axes as binary variables or as continuous spectra leaves the count unchanged: the four goal pairs are quadrants. Whether the axes are two-valued or continuous, a carrier optimized for one quadrant’s goal pair cannot simultaneously be optimized for another quadrant’s — trying to hold both ends of one goal pair level inside a single carrier is exactly the engineering pathology argued in §5.2, under either reading. The pigeonhole argument counts quadrants, not values. One further semantic note: stability here is a life-cycle property — layer content does not drift across version evolution, a matter of maintainability — not a run-time dynamical property; openness likewise means content augmentation across versions, not the system’s run-time behavior toward unseen inputs. Formalizing the four goals as properties of transition functions or information measures would move the proposition into a category the framework does not inhabit. Note: the “semi-formal” character of this appendix is a deliberate choice of limits. The two conditions of the criterion (intentional constitution, readable archive) are judged with domain knowledge and cannot be decided by pure logic; we formalize what can be formalized (direction of fit, goal incompatibility, the pigeonhole lower bound) and leave what cannot to the philosophical arguments of §3–§5 — each in its place, neither usurping the other. A.4 A Boundary Theorem: The Applicability Domain of Certificate-Anchored Calculi Definition (A.7 — archive anchoring). For a task T situated in world W, its typed context Typed(T) is legitimately grounded if and only if: (i) every type object and interface declaration of Typed(T) is extracted from a promulgated, readable archive A of W; (i) every such declaration is traceable to a clause of A. Proposition (A.4 — the boundary of certificate chains). Let C be any certificate-anchored composition calculus whose soundness theorem takes typed task contexts and certificates anchored to declared assumptions as antecedents (COIN’s Theorem 5.1 is the typical example [39]). If W satisfies ¬ -constituted(W) ∨ ¬ -archive(W), then no task in W has a legitimately grounded typed context. Backward-chaining along the antecedents: C’s soundness theorem holds at most vacuously in W; any instantiation of C in W has certificate chains whose terminal links are necessarily stipulations — declarations with no archive to endorse them. Proof. Suppose some task T in W has a legitimately grounded Typed(T). By Definition A.7(i), W would contain a promulgated readable archive, contradicting the premise. Hence no such typed context exists, and the soundness theorem’s antecedents cannot be legitimately satisfied in W; as a conditional, the theorem remains true, but vacuously. Finally, certificates in C by construction require anchored declarations; with no archive to anchor to, the terminal link of any certificate chain is a stipulation whose truth the calculus itself cannot check. Remark (A.4). Proposition A.4 does not touch the mathematics of such calculi: a conditional is vacuously true where its antecedent fails, losing nothing. What the legitimacy criterion supplies is a map of the antecedents: it says exactly which worlds can supply legitimately grounded typed contexts (those intentionally constituted and documented) and which cannot. Section 5’s four-layer framework occupies the same boundary from the other side: its knowledge layer is precisely the artifact to which engineering archives can legitimately ground. The two constructions thus meet at the criterion — one descending from symbolic composition, one ascending from physical archives. A.5 A Computational Triad: The Complexity of Coverage, Reduction, and Extraction Proposition (A.5 — open worlds resist rule coverage — a Gold-type bound). Let W be neither intentionally constituted nor documented, and let R be any finite (more generally, recursively enumerable) rule set aimed at characterizing W’s normative structure. With no promulgated source of truth, R’s adequacy can only be assessed by induction over observed instances. By Gold’s theorem [67], from positive examples alone not even regular languages are identifiable in the limit; still less can any finite rule set be confirmed to cover an open world’s normative structure. Cyc’s stagnation is therefore best read not as an engineering shortfall but as an instance of a computability-theoretic boundary: what failed was not the writing of rules but the epistemology of their confirmation. Argument. Any confirmation procedure for R in W receives only positive examples — the world presents what happened, never what should have happened; negative labels such as accident or violation records exist only where norms already exist to issue them, and Proposition A.5’s worlds are precisely those without such issuers. Gold’s theorem applies to identification from positive presentations; we assume — conservatively, for any open world — that the normative structure of an unconstituted world is at least as rich as the regular languages over its event alphabet, so no such procedure identifies that structure in the limit. Confirmation fares no better, by underdetermination rather than analogy: for any finite sample of positive presentations there are extensions of the target structure consistent with the sample on which R covers the structure, and extensions on which it does not — no finite evidence confirms exact coverage either. Observation (A.6 — decidability of failure reduction over a closed concept layer). Let the concept layer C be a finite closed set (failure-mode types, environmental factors, system transfer factors, etc., all maintained as enumerative lists — Appendix B), with inter-layer reduction proceeding along the collapse chain. Then failure localization is decidable in polynomial time: the reduction path’s depth is bounded by the number of layers d (d = 4), each layer’s branching is bounded by |C||C|, and the reduction cost is O(d⋅O(d·|C|)). Worlds without a closed concept layer offer no such guarantee — reduction may extend indefinitely, and localization is at most semi-decidable. Justification. Each reduction step moves from a concept of the current layer to a concept of an adjacent layer connected by a vertical edge; the closed vocabulary makes each layer’s candidate set finite, and the collapse chain’s length is bounded by d. The search space is therefore finite, and the procedure necessarily terminates. Observation (A.7 — the readability gradient is a complexity gradient). Extraction cost follows the way an archive was written. (i) Archives written to be read (standards, manuals, specifications): extraction degrades to parsing, O(n) in archive length. (i) Happenstance-readable archives (records, logs, accident reports from which norms must be reconstructed): extraction involves causal structure discovery, NP-hard in general [68]. (i) No archive: extraction degrades to the situation of Proposition A.5 — recursively undecidable. The historical order in which AI has penetrated the physical world — engineered facilities first, genomes and ecosystems unbroken to this day — is the empirical shadow of exactly this gradient. Justification. (i) Grammar-guided documents admit grammar-guided extraction. (i) Reconstructing latent normative structure from observational records can be modeled as learning directed-graph structure over contributing factors, whose general case is NP-hard [68]. (i) A restatement of Proposition A.5. Remark (A.5’ — the triad is one argument). Read in sequence, the triad answers three different kinds of reviewers. Proposition A.5 tells the skeptics of Section 7 why archiveless worlds do not yield to hand-written frameworks (not a cardinality accident but a Gold-type boundary). Observation A.6 tells the systems reviewer what a closed concept layer buys: decidability of localization — a property end-to-end learning currently cannot give. Observation A.7 tells the strategist why the criterion’s first successes appear where things are “written to be read” — and predicts that penetration into “happenstance-readable” and archiveless domains will follow the complexity gradient, not the domains’ enthusiasm. One final observation sharpens Observation A.6 from the other side — a model that commands only distributional similarity can flag the unlikely, but cannot decide whether an input violates a norm or instantiates a novel but legitimate behavior; we give the full argument, and the operational stakes of the distinction, in §5.4. Appendix B. The Failure-Mode Type Vocabulary (Excerpt) This appendix gives an excerpt of the failure-mode type vocabulary in the concept layer, as engineering-inventory evidence for the claim “types are enumerable” (§5.3). The skeleton’s four-layer shape is unfolded in §5.3; here we add only the implementation facts the main text does not carry. First, the concept layer’s internal structure: an industry’s object system reduces to six object classes (physical objects, spatial objects, asset objects, operational objects, organizational objects, virtual objects), invariant across industries, augmented but never altered as coverage expands; hundreds of domain concepts are all attached by inheritance under the syntax layer’s nouns — the inter-layer vertical edges are literally inheritance/instantiation relations in the implementation. Second, what passes between layers is neither discrete events nor continuous physical quantities but graded state semantics: 0 normal / 0.25 adverse trend / 0.5 symptom / 1 failure — a decision-facing rather than physics-facing notation, rewritable by observation or by inference, its reduction proceeding along the inter-layer vertical edges. Third, for newly appearing high-value artifacts, modeling follows the top-down path: design first, artifact after — their construction and commissioning phases have typically already exercised the vast majority of failure modes — the knowledge layer’s initial content is likewise a gift of the design archive. Proprietary implementation details of the framework are commercial information and are omitted here. B.1 Physical-System Failure Mechanisms (10 classes) Modes of action by which systems, equipment, or components suffer interruption, degradation, or destabilization of designed function due to anomalies of structure, function, flow, or control capability. # Type Definition 1 structural rupture failure fracture or breakage of physical structure due to mechanical stress, fatigue, or corrosion 2 seal leakage failure medium leakage due to seal aging, poor assembly, or pressure anomaly 3 parameter drift failure sensing or control parameters deviating from normal range, causing control inaccuracy 4 function interruption failure complete loss of core function; design intent can no longer be executed 5 degradation and wear failure progressive performance decline or mechanical wear from long-term operation 6 passage blockage failure flow passages, piping, or channels obstructed by fouling, foreign matter, or deformation 7 system oscillation failure periodic abnormal fluctuation due to control-loop instability or mechanical resonance 8 constraint violation failure operating parameters exceeding design boundaries or safety thresholds 9 instability amplification failure state deviation continuously amplified by positive feedback 10 connection breakage failure interruption or contact degradation of physical, electrical, or signal connections B.2 Medium-State Failure Mechanisms (9 classes) State-change processes in which the composition, physical properties, or phase of a system’s internal medium deviates or degrades, reducing its functional support capacity. # Type Definition 1 composition contamination failure unintended impurities mixed in; chemical purity or functional properties impaired 2 concentration deviation failure active-ingredient concentration deviating from design value 3 phase anomaly failure unintended phase change (vaporization, crystallization, solidification), losing original flow or heat-transfer properties 4 property degradation failure degradation of viscosity, thermal conductivity, surface tension, or other physical properties 5 chemical deactivation failure active chemical components losing reactivity (catalyst poisoning, corrosion-inhibitor failure) 6 structural decomposition failure decomposition of macromolecular or composite structures (polymer degradation, emulsion breaking) 7 adsorption saturation failure adsorptive media reaching capacity limits 8 consumable exhaustion failure consumable media exhausted (desiccant saturated, filter capacity used up) 9 medium destabilization failure stratification, sedimentation, or agglomeration; loss of homogeneous stability B.3 System Transfer Factors (4 classes) Cross-system supply anomalies that equipment/spaces receive from beyond the object boundary. Classes are divided on a single dimension by the physical category of the transferred: energy / matter / information / mechanical. An earlier two-axis scheme (transfer type × transfer state: supply interruption / supply anomaly) has been collapsed — interruption is treated as the extreme value of deviation, with severity carried by the four-grade failureDegreeProfile; one axis fewer, no omission space. A classification decision tree assigns every instance uniquely to one of the four classes via three binary questions. # Type Definition and examples 1 matter tangible transfer objects with mass: fluid media (water, gas, steam, fuel oil), solid materials, and persons and goods whose transport is the service purpose — chilled-water flow, gas flow, fresh-air volume 2 energy energy forms transferred via separable carriers or without carriers: electric power, cooling, heat, pressure energy — supply voltage, cooling capacity, steam heat 3 information semantic content whose value lies in “what it expresses,” independent of the carrying physical quantity: control commands, feedback signals, measurement data — setpoint correctness, measurement-point data accuracy, communication-link availability 4 mechanical force, motion, or load transmitted through inseparable rigid/semi-rigid structures: active transmission, static load-bearing, structural vibration transmission — transmission torque, structural bearing capacity, transmission alignment B.4 Environmental Factors (9 classes) Types of physical interference in the equipment’s operating environment or boundary conditions that exceed design tolerances. # Type Definition 1 ambient temperature anomaly temperature deviating from the designed operating range; thermal balance destabilized 2 ambient humidity anomaly humidity or liquid water vapor beyond design boundaries; insulation degradation, condensation, or corrosion risk 3 mechanical environment anomaly vibration, shock, pressure waves, or other mechanical loads beyond design endurance 4 electromagnetic interference anomaly external electromagnetic fields beyond the equipment’s immunity; signal distortion or control anomalies 5 spatial constraint anomaly restricted space, airflow, or heat-dissipation paths; degraded operating boundary conditions 6 medium intrusion anomaly water, gas, dust, or other media entering through non-designed paths, altering physical boundary conditions 7 foreign-object intrusion anomaly small animals, tools, metal fragments, or other discrete entities entering the operating space 8 chemical environment anomaly corrosive gases, liquids, or active substances beyond equipment tolerance 9 radiation environment anomaly ionizing, non-ionizing, or thermal radiation beyond design protection capability Appendix C. Framework Demonstrations Preamble: The Existence Claim Claim — prior cognitive frameworks satisfying the structure above are in actual operation in industrial facility operations across five major domains, covering civil public buildings (shopping malls, hotels, offices, hospitals), industrial manufacturing (semiconductors, lithium batteries, electronics), major scientific installations (a nuclear-fusion experimental device, communication satellites), municipal infrastructure (district heat networks), and data centers (the dwelling place of artificial intelligence itself); in cross-industry migration, the concept layer migrates unchanged — each new industry “only adds blocks, never changes the shape” — and the marginal modeling cost decreases with the number of industries, and, as stated in §5.3, onboarding a new industry has required only instantiating known failure types, never inventing new ones. Commitment — the framework’s cross-domain invariance is a publicly testable falsifiable proposition (Section 6, P3): if the onboarding of any new industry requires modifying the concept layer, this paper’s propositions are thereby damaged. The Demonstration Register: Selection, Disclosure Status, and Commitments The framework’s demonstrations are organized as a register of four cases in two tiers, the two tiers answering two different questions. The first tier demonstrates that a legitimate prior framework can reason: C.1, the Curiosity Mars rover’s drill-feed-mechanism anomaly (Sol 1536, December 2016; NASA public mission materials [71]) — systems without piping, and at the same time the limiting case of prior-framework intelligence: a single, physically inaccessible instance for which no failure training set can ever exist, so that the design archive is the only asset there is; C.2, a chiller plant — piped systems, the everyday counterpart of C.1. Each first-tier report walks the full path: component decomposition at the instance layer, failure-mode instance definitions as instantiations of Appendix B’s types at the concept layer, and one complete reasoning chain — how a field alarm collapses along the vertical edges to its cause, and how a preventive decision follows. The second tier demonstrates that the concept layer does not drift across industries: C.3, a factory pure-water system, and C.4, a city heat network. These two reports carry failure-mode instantiation statements only — which types of Appendix B each industry’s observed failure modes instantiate — because that is all their question requires: the evidence for the invariance prediction of Section 6 (P3) is that a new industry instantiates known types and invents none. Disclosure status, stated plainly. The four case reports are not yet complete at this version, and we say so on the record rather than gesture at them. Two of the register’s inputs are already fixed and public: the register itself (what will be demonstrated, by which tier, and why these four), and the event record of C.1, which lives entirely in NASA’s public mission materials [71]. The reports themselves are withheld as one set, pending de-sensitization and patent applications, per the “claims + commitments” protocol of §1. The reason lies in what a case report of this paper is: not a re-narration of events but a re-enactment of them through the framework — the event record is the background; what the report demonstrates is how a prior framework built to this paper’s architecture expresses the instance and reasons over it — so that any full report, C.1’s included, necessarily discloses the layering and reasoning mechanism in operation. All four reports will be added together in a revised version of this paper. Why the register is not a promissory note. A demonstration register without its reports would normally ask the reader for trust; this one does not. What the register carries is already checkable: the selection rationale is stated above; the C.1 event record is public; and the register’s theory-bearing content — that the four-layer skeleton carries a complete reasoning chain from alarm to located cause (first tier), and that the concept layer of Appendix B covers each new industry by instantiation, not by invention (second tier) — is staked as the falsifiable commitments of Section 6, which any reader can attack on the public engineering record without waiting for any of our reports; the second tier’s content is prediction P3 itself. The register fixes, in advance, which observations would damage the framework; that is its function here. References [1] D.B. Lenat, CYC: A large-scale investment in knowledge infrastructure, Communications of the ACM 38 (11) (1995) 33–38. [2] D. Silver, T. Hubert, J. Schrittwieser, et al., A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play, Science 362 (6419) (2018) 1140–1144. [3] M. Chen, J. Tworek, H. Jun, et al., Evaluating large language models trained on code, arXiv:2107.03374, 2021. [4] OpenAI, GPT-4 technical report, arXiv:2303.08774, 2023. [5] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, B. Ommer, High-resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, p. 10684–10695; arXiv:2112.10752. [6] H. Moravec, Mind Children: The Future of Robot and Human Intelligence, Harvard University Press, Cambridge, MA, 1988. [7] S. Harnad, The symbol grounding problem, Physica D 42 (1–3) (1990) 335–346. [8] D. Ha, J. Schmidhuber, Recurrent world models facilitate policy evolution, in: Advances in Neural Information Processing Systems 31 (NeurIPS), 2018, p. 2450–2462; arXiv:1803.10122. [9] Y. LeCun, A path towards autonomous machine intelligence, OpenReview Preprint, 2022. [10] J. Bruce, M.D. Dennis, A. Edwards, et al., Genie: Generative interactive environments, in: Proceedings of the 41st International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research 235 (2024) 4603–4623; arXiv:2402.15391. [11] NVIDIA, Cosmos world foundation model platform for physical AI, arXiv:2501.03575, 2025. [12] J. von Uexkül, G. Kriszat, Streifzüge durch die Umwelten von Tieren und Menschen: Ein Bilderbuch unsichtbarer Welten, Verständliche Wissenschaft, vol. 21, Springer, Berlin, 1934, https://doi.org/10.1007/978-3-642-98976-6. [13] D.C. Dennett, The Intentional Stance, MIT Press, Cambridge, MA, 1987. [14] T.R. Gruber, A translation approach to portable ontology specifications, Knowledge Acquisition 5 (2) (1993) 199–220, https://doi.org/10.1006/knac.1993.1008. [15] N. Guarino, Formal ontology and information systems, in: N. Guarino (Ed.), Formal Ontology in Information Systems, IOS Press, Amsterdam, 1998, p. 3–15. [16] H.A. Simon, The Sciences of the Artificial, MIT Press, Cambridge, MA, 1969. [17] K.R. Popper, Objective Knowledge: An Evolutionary Approach, Oxford University Press, Oxford, 1972. [18] N. Hartmann, Der Aufbau der realen Welt: Grundriss der allgemeinen Kategorienlehre, De Gruyter, Berlin, 1940. [19] F.-Y. Wang, Parallel system methods for management and control of complex systems, Control and Decision 19 (5) (2004) 485–489, 514, https://doi.org/10.3321/j.issn:1001-0920.2004.05.002. [20] M. Tomasello, Constructing a Language: A Usage-Based Theory of Language Acquisition, Harvard University Press, Cambridge, MA, 2003. [21] J. Bybee, Language, Usage and Cognition, Cambridge University Press, Cambridge, 2010. [22] M. Raissi, P. Perdikaris, G.E. Karniadakis, Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations, Journal of Computational Physics 378 (2019) 686–707. [23] P. Kroes, A. Meijers, The dual nature of technical artefacts, Studies in History and Philosophy of Science 37 (1) (2006) 1–4. [24] J.R. Searle, The Construction of Social Reality, Free Press, New York, 1995. [25] A.W. Senior, R. Evans, J. Jumper, et al., Improved protein structure prediction using potentials from deep learning, Nature 577 (2020) 706–710. [26] J. Jumper, R. Evans, A. Pritzel, et al., Highly accurate protein structure prediction with AlphaFold, Nature 596 (2021) 583–589. [27] M. Polanyi, The Tacit Dimension, Doubleday, Garden City, NY, 1966. [28] G.E.M. Anscombe, Intention, Basil Blackwell, Oxford, 1957. [29] J.R. Searle, Intentionality: An Essay in the Philosophy of Mind, Cambridge University Press, Cambridge, 1983. [30] G.A. Carpenter, S. Grossberg, A massively parallel architecture for a self-organizing neural pattern recognition machine, Computer Vision, Graphics, and Image Processing 37 (1) (1987) 54–115. [31] F. Baader, D. Calvanese, D. McGuinness, D. Nardi, P.F. Patel-Schneider (Eds.), The Description Logic Handbook: Theory, Implementation and Applications, Cambridge University Press, Cambridge, 2003. [32] J. Rasmussen, The role of hierarchical knowledge representation in decisionmaking and system management, IEEE Transactions on Systems, Man, and Cybernetics SMC-15 (2) (1985) 234–243. [33] K.J. Vicente, J. Rasmussen, Ecological interface design: Theoretical foundations, IEEE Transactions on Systems, Man, and Cybernetics 22 (4) (1992) 589–606. [34] J. Hirtz, R.B. Stone, D.A. McAdams, S. Szykman, K.L. Wood, A functional basis for engineering design: Reconciling and evolving previous efforts, Research in Engineering Design 13 (2) (2002) 65–82, https://doi.org/10.1007/s00163-001-0008-3. [35] SAE International, Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles, SAE Standard J3016, 2021. [36] International Organization for Standardization, Road Vehicles — Test Scenarios for Automated Driving Systems — Specification for Operational Design Domain, ISO 34503:2023, Geneva, 2023. [37] T. Charmet, V. Cherfaoui, J. Ibanez-Guzman, A. Armand, Overview of the operational design domain monitoring for safe intelligent vehicle navigation, in: Proceedings of the 26th IEEE International Conference on Intelligent Transportation Systems (ITSC), 2023, p. 5363–5370, https://doi.org/10.1109/ITSC57777.2023.10421823. [38] ASAM e.V., ASAM OpenODD Concept Paper: Concepts for a Machine-Interpretable Description of the Operational Design Domain, Version 1.0, ASAM e.V., Höhenkirchen-Siegertsbrunn, 2021. [39] Y. Yuan, A.C.-C. Yao, Calculus of intelligence: A topos-monadic framework for agentic workflows, iFuture 1 (2026) 9710001, https://doi.org/10.26599/IF.2026.9710001. [40] Palantir Technologies, Foundry Documentation: Ontology Overview, https://w.palantir.com/docs/foundry/ontology/overview, accessed 2 August 2026. [41] A. Brohan, N. Brown, J. Carbajal, et al., RT-2: Vision-language-action models transfer web knowledge to robotic control, in: Proceedings of the 7th Conference on Robot Learning (CoRL), Proceedings of Machine Learning Research 229 (2023) 2165–2183; arXiv:2307.15818. [42] K. Black, N. Brown, D. Driess, et al., π0 _0: A vision-language-action flow model for general robot control, in: Proceedings of Robotics: Science and Systems XXI (RSS), 2025; arXiv:2410.24164. [43] I. Lakatos, The Methodology of Scientific Research Programmes, Cambridge University Press, Cambridge, 1978. [44] C. Perrow, Normal Accidents: Living with High-Risk Technologies, Basic Books, New York, 1984. [45] N. Guarino, C.A. Welty, An overview of OntoClean, in: S. Staab, R. Studer (Eds.), Handbook on Ontologies, Springer, Berlin, 2004, p. 151–171, https://doi.org/10.1007/978-3-540-24750-0_8. [46] R.J. Brachman, H.J. Levesque, Knowledge Representation and Reasoning, Morgan Kaufmann, San Francisco, 2004. [47] J. McCarthy, Circumscription—A form of non-monotonic reasoning, Artificial Intelligence 13 (1–2) (1980) 27–39. [48] W3C OWL Working Group, OWL 2 Web Ontology Language Document Overview, 2nd ed., W3C Recommendation, 2012. [49] I. Horrocks, P.F. Patel-Schneider, H. Boley, S. Tabet, B. Grosof, M. Dean, SWRL: A semantic web rule language combining OWL and RuleML, W3C Member Submission, 2004. [50] G.H. von Wright, Deontic logic, Mind 60 (237) (1951) 1–15. [51] M. Esteva, J.A. Rodríguez-Aguilar, C. Sierra, P. García, J.L. Arcos, On the formal specification of electronic institutions, in: Agent Mediated Electronic Commerce, Springer, Berlin, 2001, p. 126–147. [52] A.S. Rao, M.P. Georgeff, BDI agents: From theory to practice, in: Proceedings of the First International Conference on Multi-Agent Systems (ICMAS’95), 1995, p. 312–319. [53] M. Dastani, 2APL: A practical agent programming language, Autonomous Agents and Multi-Agent Systems 16 (3) (2008) 214–248. [54] G. Boella, L. van der Torre, Regulative and constitutive norms in normative multiagent systems, in: Proceedings of the Ninth International Conference on Principles of Knowledge Representation and Reasoning (KR’04), 2004, p. 255–266. [55] A. Singhal, Introducing the Knowledge Graph: Things, not strings, Google Official Blog, 2012. [56] B. Balaji, A. Bhattacharya, G. Fierro, et al., Brick: Towards a unified metadata schema for buildings, in: Proceedings of the 3rd ACM International Conference on Systems for Energy-Efficient Built Environments (BuildSys’16), 2016, p. 41–50. [57] Project Haystack, Project Haystack: Semantic modelling for device and equipment data, https://project-haystack.org, accessed 3 August 2026. [58] International Organization for Standardization, Road Vehicles — Safety of the Intended Functionality, ISO 21448:2022, Geneva, 2022. [59] A. d’Avila Garcez, L.C. Lamb, Neurosymbolic AI: The 3rd wave, Artificial Intelligence Review 56 (2023) 12387–12406, https://doi.org/10.1007/s10462-023-10448-w. [60] J. Mao, C. Gan, P. Kohli, J.B. Tenenbaum, J. Wu, The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision, in: International Conference on Learning Representations (ICLR), 2019. [61] R. Manhaeve, S. Dumančić, A. Kimmig, T. Demeester, L. De Raedt, DeepProbLog: Neural probabilistic logic programming, in: Advances in Neural Information Processing Systems 31 (NeurIPS), 2018. [62] S. Badreddine, A. d’Avila Garcez, L. Serafini, M. Spranger, Logic tensor networks, Artificial Intelligence 303 (2022) 103649. [63] S. Yao, J. Zhao, D. Yu, et al., ReAct: Synergizing reasoning and acting in language models, in: International Conference on Learning Representations (ICLR), 2023. [64] T. Schick, J. Dwivedi-Yu, R. Dessì, et al., Toolformer: Language models can teach themselves to use tools, in: Advances in Neural Information Processing Systems 36 (NeurIPS), 2023; arXiv:2302.04761. [65] C. Packer, V. Fang, S.G. Patil, K. Lin, S. Wooders, I. Gonzalez, MemGPT: Towards LLMs as operating systems, arXiv:2310.08560, 2023. [66] Q. Wu, G. Bansal, J. Zhang, et al., AutoGen: Enabling next-gen LLM applications via multi-agent conversation, arXiv:2308.08155, 2023. [67] E.M. Gold, Language identification in the limit, Information and Control 10 (5) (1967) 447–474. [68] D.M. Chickering, Learning Bayesian networks is NP-complete, in: D. Fisher, H.-J. Lenz (Eds.), Learning from Data: Artificial Intelligence and Statistics V, Springer, New York, 1996, p. 121–130. [69] J. Jiang, Form follows purpose: Why the artificial physical world needs brains, not bodies, in preparation, 2026. [70] J. Jiang, The missing carrier: On the necessity and conditional sufficiency of LLMs for machine intelligence in the artificial physical world, in preparation, 2026. [71] NASA Jet Propulsion Laboratory, Mars Science Laboratory mission status reports on the Curiosity rover drill feed mechanism anomaly, December 2016, https://mars.nasa.gov/msl/. [72] P. Kroes, Technical Artefacts: Creations of Mind and Matter: A Philosophy of Engineering Design, Springer, Dordrecht, 2012. [73] L. Floridi, The method of levels of abstraction, Minds and Machines 18 (3) (2008) 303–329. [74] T.P. Kelly, Arguing Safety: A Systematic Approach to Managing Safety Cases, PhD thesis, University of York, 1998. [75] R. Hawkins, C. Paterson, C. Picardi, et al., Guidance on the assurance of machine learning in autonomous systems (AMLAS), University of York, 2021.