Paper deep dive
Time, Identity and Consciousness in Language Model Agents
Elija Perrier, Michael Timothy Bennett
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 92%
Last extracted: 3/13/2026, 12:57:57 AM
Summary
The paper introduces a formal framework for evaluating identity in Language Model Agents (LMAs) by applying Stack Theory's temporal semantics. It distinguishes between 'ingredient-wise occurrence' (where identity components appear within a window) and 'co-instantiation' (where they are active simultaneously at a decision step). The authors demonstrate that LMAs can exhibit stable self-report behavior while failing to co-instantiate identity constraints, leading to a 'temporal gap' that undermines reliability and consciousness attribution. They provide a toolkit of persistence scores and identity metrics to measure this gap.
Entities (5)
Relation Signals (3)
Stack Theory → defines → Temporal Gap
confidence 95% · We apply Stack Theory's temporal gap to scaffold trajectories.
Language Model Agents → exhibits → Temporal Gap
confidence 90% · An LMA can satisfy ingredient-wise identity checks for multiple identity ingredients across a window while still failing to ever instantiate the full identity conjunction.
Arpeggio → measures → Identity
confidence 85% · We use these postulates to measure the window-level occurrence and co-instantiation conditions.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Machine consciousness evaluations mostly see behavior. For language model agents that behavior is language and tool use. That lets an agent say the right things about itself even when the constraints that should make those statements matter are not jointly present at decision time. We apply Stack Theory's temporal gap to scaffold trajectories. This separates ingredient-wise occurrence within an evaluation window from co-instantiation at a single objective step. We then instantiate Stack Theory's Arpeggio and Chord postulates on grounded identity statements. This yields two persistence scores that can be computed from instrumented scaffold traces. We connect these scores to five operational identity metrics and map common scaffolds into an identity morphospace that exposes predictable tradeoffs. The result is a conservative toolkit for identity evaluation. It separates talking like a stable self from being organized like one.
Tags
Links
- Source: https://arxiv.org/abs/2603.09043v1
- Canonical: https://arxiv.org/abs/2603.09043v1
Trouble viewing inline? Open PDF directly →
Full Text
57,364 characters extracted from source content.
Expand or collapse full text
Time, Identity and Consciousness in Language Model Agents * Elija Perrier 1 † , Michael Timothy Bennett 2† 1 Centre for Quantum Software and Information, UTS, Sydney 2 Australian National University, Canberra elija.perrier@gmail.com m@michaeltimothybennett.com Abstract Machine consciousness evaluations mostly see behavior. For language model agents that behavior is language and tool use. That lets an agent say the right things about itself even when the constraints that should make those statements mat- ter are not jointly present at decision time. We apply Stack Theory’s temporal gap to scaffold trajectories. This separates ingredient-wise occurrence within an evaluation window from co-instantiation at a single objective step. We then instantiate Stack Theory’s Arpeggio and Chord postulates on grounded identity statements. This yields two persistence scores that can be computed from instrumented scaffold traces. We con- nect these scores to five operational identity metrics and map common scaffolds into an identity morphospace that exposes predictable tradeoffs. The result is a conservative toolkit for identity evaluation. It separates talking like a stable self from being organized like one. 1 Introduction Machine consciousness research is short on direct evidence. For artificial agents the safest evidence we can collect is behavioral. For language model agents (LMAs) most of that behavior is language, tool use, and the traces they leave in external memory. This creates a trap. A system can talk like it has a stable self while the underlying identity constraints that should govern its actions are never jointly active at decision time. A scaffold can make identity ingredients retrievable with- out making them jointly active at action time. For exam- ple, an agent may reliably restate its name, role, and safety constraints when queried about each in isolation. Yet when it must choose an action, those ingredients can fail to co- instantiate in the decision state. That is how an agent can talk in character while acting out of character. This paper applies Stack Theory’s temporal gap to agent identity in LMAs (Bennett 2025, 2026a). The temporal gap is the logical gap between ingredient-wise occurrence within * Accepted at AAAI 2026 Spring Symposium - Machine Con- sciousness: Integrating Theory, Technology, and Philosophy † These authors contributed equally. Copyright © 2026, Association for the Advancement of Artificial Intelligence (w.aaai.org). All rights reserved. a window and co-instantiation at a single objective step. Oc- currence means each identity ingredient is active somewhere in the window. Co-instantiation means there is a single objec- tive step where the full identity conjunction is active. Many common scaffolds can achieve occurrence without reliably achieving co-instantiation. That is why an agent can pass recall-based identity tests and still act out of character when the decision actually matters. 1.1 The Challenge of LMA Identity As AI systems become increasingly autonomous, agent iden- tity becomes crucial to reliability, safety, and utility. Identity asks whether a system remains the same agent over time and across contexts. LMAs present a unique challenge. They situ- ate an LLM inside an agentic scaffold of prompts, memory modules, retrieval, and tool APIs to enable planning, reason- ing, and action (Kapoor et al. 2024; Liu et al. 2023; Wu 2024). Yet the core LLM is stateless at inference. It only sees the current input. Any persistent identity must be reconstructed from external traces. This paper answers two precise questions. What does it mean for an LMA to preserve its identity over time? Under what formal conditions is that even possible? The problem. Existing discussions of LMA identity are informal. Terms like statelessness, persistence, and identity drift are used without precise definitions. This imprecision hides how an identity component can occur somewhere in the recent interaction history without constraining the current decision. An agent might separately state its name, role, con- straints, and goals across different turns without ever having a time slice where the full identity conjunction is simultane- ously active. Our approach. We treat the scaffold state space as the environment and apply Stack Theory’s window semantics to scaffold trajectories (Bennett 2026a). We then restate the temporal gap result in this setting. In particular, the within window diamond lift does not distribute over conjunction (Theorem 3.10). This separates ingredient-wise recall from operative identity. We then use Stack Theory’s Arpeggio and Chord postulates as an interpretive lens for identity in the machine consciousness setting (Bennett 2026a). We use arXiv:2603.09043v1 [cs.AI] 10 Mar 2026 these postulates to measure the window-level occurrence and co-instantiation conditions that Arpeggio and Chord appeal to. Why this matters.Identity affects three questions that the machine consciousness workshop explicitly cares about. It affects measurement, implementation, and ethics. •For evaluation. Benchmarks that test whether agents can recall identity facts may give false confidence. An agent that passes recall tests can still fail to act according to its identity because recall does not imply co-instantiation. •For design. Retrieval and memory systems can improve ingredient availability, but can also fragment identity by surfacing competing fragments. This is a predictable con- sequence of the temporal gap. •For safety and moral status. Safety constraints must be co-instantiated with goals during action selection. Moral status debates also become harder when the target of at- tribution is not stable across time. If you cannot say what the agent is at a moment, you cannot cleanly ask whether that moment is conscious. Relevance to machine consciousness. Many conscious- ness proposals require some form of integration that binds the contents of a moment into a single subject, even if they disagree about what that integration is (Bennett 2025; Baars 1988; Dehaene and Naccache 2001; Tononi 2004; Metzinger 2003). Some proposed indicators for AI consciousness there- fore lean on behavior that looks like a stable self model. This includes self-report, memory, and narrative continuity (Bennett 2023a, 2025, 2026b, 2023b). Our results isolate a specific failure mode for such indicators. A system can look stable under self-report while failing to ever co-instantiate the grounded identity conjunction that would make that stability operative. Contributions. This paper makes the following contribu- tions. 1.Temporal semantics for LMA identity. We apply win- dowing maps, occurrence predicates, and co-instantiation conditions that precisely characterise when identity is preserved in LMAs. 2.Arpeggio and Chord applied to identity. We restate Stack Theory’s Arpeggio and Chord postulates and show how their Occur versus CoInst consequents become mea- surable identity criteria in LMA scaffolds. 3. Compositional grounding. We formalise the layered structure of identity from implementation variables (Layer 0) through functional commitments (Layer 1) to narrative self model (Layer 2). 4.Identity morphospace. Drawing on cognition science (Solé et al. 2026), we organize identity metrics into a structured space and identify architectural tradeoffs and predicted voids. 5. Derived identity metrics. We show how five operational metrics emerge from the temporal theory. The metrics are Identifiability, Continuity, Consistency, Persistence, and Recovery. We also prove simple bounds on identity preservation under common scaffold configurations and explain counterintuitive effects such as retrieval reducing co-instantiation (see Appen- dices via (Perrier 2026)). How this fits into the machine consciousness discourse. Machine consciousness discourse links theory, measurement, implementation, and ethics. This paper is organized around that bridge. Sections 3 and 4 are the theoretical backbone. Section 5 gives a measurement recipe that can be run on real systems. The Discussion section connects the resulting failure modes to consciousness attribution and to the ethics of deploying agents that can convincingly self-narrate while failing to bind their constraints in action. 2 Formal Scaffold Model Before the temporal semantics, we introduce a minimal for- mal model of LMA scaffolds. We treat the scaffold state spaceSas the Stack Theory environment, and we treat each grounded identity ingredient as a program g 0 i ⊆ S. This lets us apply the Stack Theory definitions of conjunction, win- dowing, occurrence, and co-instantiation directly. This model captures the essential components that determine what infor- mation is available to the LLM at any given moment, which in turn determines what aspects of identity can be “active” during decision-making. We focus at the scaffold level because it is where iden- tity becomes enforceable. It is also where identity becomes measurable, because we can instrument which grounded in- gredients are active and when. The key insight is that an LMA’s operative identity at any moment is whatever is in the token sequence the LLM ac- tually processes during inference. If an identity ingredient is not effectively present there, then it cannot constrain the next action. Our model makes the main context sources ex- plicit. They include conversation history, external memory, retrieved documents, and policy flags. Definition 2.1 (Scaffold architecture). A scaffold architec- ture is a tupleA = (Σ,K,V,Q,D,R,n π ,|C| max ). • Σ is the token alphabet. • K and V are the key and value sets for external memory. • Q is the query space. • D is the document corpus available to retrieval. • R : Q→ 2 D is the retrieval function. • n π ∈N >0 is the number of binary policy flags. • |C| max ∈N >0 is the context capacity measured in tokens. Definition 2.2 (Scaffold state). Fix an architectureA. A scaffold state is a tuple s = (C,M,π,D retrieved ) where • C ∈ Σ ∗ is the current context window and|C|≤|C| max . • M : K ⇀ V is the current memory store contents. • π ∈0, 1 n π is the current policy flag vector. • D retrieved ⊆ Dis the set of retrieved documents currently injected. We writes.C,s.M,s.π,s.Dfor the components. LetSbe the set of all scaffold states consistent withA. Definition 2.3 (Scaffold transition). A scaffold transition functionδ : S × A → Smaps current state and action to next state. Actions A include • infer(q) is LLM inference with query q. It updates C. • retrieve(q)is retrieval augmented generation. It up- dates D retrieved . • store(k,v) is a memory write. It updates M . • tool(t,args)is a tool call. It may update any compo- nent. Definition 2.4 (Ingredient activation). An identity ingredient g 0 i is active in states(writtens|= g 0 i ) iff the required imple- mentation level condition is present insin a way that can affect the next inference. Concretely this means the follow- ing. • Ifg 0 i is a context condition, then the required tokens ap- pear in s.C. •Ifg 0 i is a memory condition, then the required key value pairs exist in s.M . • Ifg 0 i is a policy condition, then the required flags are set in s.π. •Ifg 0 i is a retrieval condition, then the required document is in s.D. The full grounded identityg 0 = g 0 1 ∧·∧ g 0 k is active ins iff all ingredients are active. An identity ingredient is not “active” simply because it exists somewhere in the system’s storage. It is active only if the relevant information is present in the current state in a way that can influence the LLM’s output. This is the formal counterpart to our intuitive distinction between “retrievable” and “decision-guiding.” Example 2.5 (Activation in Practice). Consider an agent with identity “helpful assistant focused on privacy.” The ingredient g 0 privacy requires that privacy-related tokens appear in context. This ingredient is: • Active if “privacy” appears in the system prompt currently ins.C, or if a privacy policy document is ins.D, or if a privacy flag is set in s.π. •Not active if privacy information exists only in the mem- ory stores.Mbut was not retrieved into context for this inference. The ingredient may be stored (available for future retrieval) without being active (influencing current behavior). This distinction is one concrete instance of the temporal gap. This model is minimal but sufficient to formalize our ar- chitectural theorems. 3 Temporal Semantics for Agent Identity We apply Stack Theory’s temporal semantics to the LMA setting (Bennett 2026a). The key distinction is between iden- tity ingredients that occur somewhere in a recent window and identity ingredients that are co-instantiated at a single objective step. 3.1 Objective time, layer time, and windowing Definition 3.1 (Agent trajectory). An agent trajectory is a functionτ :N→ Sthat maps each objective time stepu∈N to a scaffold state s u = τ (u). Objective time indexes the actual computational micro steps. These are LLM calls, tool invocations, retrieval opera- tions, and memory updates. Users reason at a coarser time scale. They ask identity questions at the level of turns, tasks, and episodes. We model this coarser time scale by indexing windows over objective time. Definition 3.2 (Windowing map). Fix a horizon∆∈Nand a strides ∈N >0 . The windowing mapW ∆,s sends a layer time index t∈N to a windowed trajectory segment W ∆,s (t) = τ (st),τ (st + 1),...,τ (st + ∆) .(1) If∆ = 0this is a one step window andW 0,s (t) = (τ (st)). We writeσ ∆,s (t)for this windowed segment. When∆ands are clear from context we write W (t) and σ(t). This is the same window construction used in Stack Theory (Bennett 2026a). The horizon∆controls how forgiving the evaluation is. A larger∆allows identity ingredients to be spread across more objective steps. A smaller∆demands tighter temporal coherence. 3.2 Identity statements and grounding An agent’s identity is typically described at a high level. For example, an agent might be described as a privacy-focused data analyst. Grounding makes explicit what this means in terms of the underlying scaffold state. Definition 3.3 (Identity statement). An identity statementl m at layer m is a conjunction of identity predicates l m = p m 1 ∧ p m 2 ∧·∧ p m n (2) where eachp m i is an atomic identity predicate such as name, role, goal, or constraint. Definition 3.4 (Grounding operation). The grounding oper- ationGround 0←m : L m → L 0 maps identity statements at layer m to implementation level requirements at layer 0 Ground 0←m (p m 1 ∧·∧ p m n ) = g 0 1 ∧·∧ g 0 k (3) where eachg 0 j is a condition on implementation variables such as system prompt tokens, memory slot contents, tool outputs, controller flags, or policy parameters. Grounding turns abstract identity claims into concrete com- putational conditions. Name equals Alice can ground to a requirement that the token Alice appears in the system prompt or in a pinned context region. Constraint equals privacy can ground to a requirement that a privacy policy is present in context, or that a privacy flag is set, or that a tool is disabled. Definition 3.5 (Grounded identity). Given an identity state- ment l m , its grounded identity is g 0 = Ground 0←m (l m ).(4) 3.3 Occurrence versus co-instantiation Letg 0 = g 0 1 ∧·∧ g 0 k be a grounded identity conjunction. Relative to a windowW (t)there are two ways to ask whether identity is present. Definition 3.6 (Window satisfaction). Letσ(t) = W (t) = (s st ,...,s st+∆ ) be the window at layer time t. • Occur W (g 0 ,τ,t) holds iff for each conjunctg 0 i there ex- ists an indexj i ∈0,..., ∆such thats st+j i |= g 0 i . Each identity ingredient occurs somewhere in the window. • CoInst W (g 0 ,τ,t)holds iff there exists an indexj ∈ 0,..., ∆ such thats st+j |= g 0 . All identity ingredi- ents are co-instantiated at a single objective step inside the window. Occurrence is ingredient-wise coverage. Co-instantiation is joint availability. Co-instantiation implies occurrence, but occurrence does not imply co-instantiation. Remark3.7.IfCoInst W (g 0 ,τ,t)holdsthen Occur W (g 0 ,τ,t) holds. Proof.If all conjuncts hold at the same objective step in the window, then each conjunct also holds somewhere in the window. 3.4 Temporal lifts and the temporal gap Stack Theory expresses window-level predicates using tem- poral lifts (Bennett 2026a). For a program or predicatepover scaffold states, define the within window diamond lift. Definition 3.8 (Existential temporal lift). Letpbe a predicate over scaffold states. Define ♢ ∆ p holds at layer time t iff∃j ∈0,..., ∆ such that s st+j |= p. Remark 3.9 (Occurrence and co-instantiation as lifts). Let g 0 = g 0 1 ∧·∧g 0 k . ThenOccur W (g 0 ,τ,t)holds iff♢ ∆ g 0 1 ∧ ·∧♢ ∆ g 0 k holds at layer timet. AndCoInst W (g 0 ,τ,t)holds iff♢ ∆ (g 0 )holds at layer timet. So the temporal gap is exactly the difference between lifting ingredients separately and lifting the whole conjunction at once. The central subtlety is that♢ ∆ does not distribute over conjunction. This is a standard fact in modal logic. Here it becomes a concrete failure mode for LMA identity. Theorem 3.10 (Non-commutation with conjunction). For predicates p and q over scaffold states, ♢ ∆ (p∧ q) ⇒♢ ∆ p∧♢ ∆ q(5) but the converse implication fails in general. Equivalently, ♢ ∆ (p∧ q)̸⇔♢ ∆ p∧♢ ∆ q. Proof. Ifp∧ qholds at some objective step in the window, thenpholds at that step andqholds at that step. So♢ ∆ (p∧q) implies both♢ ∆ p and♢ ∆ q. For the converse, fix a two step window with∆ = 1. Let τ (st) |= p∧¬qandτ (st + 1) |= q∧¬p. Then♢ ∆ pholds and♢ ∆ qholds, but there is no step wherep∧ qholds. So ♢ ∆ (p∧ q) fails. Corollary 3.11 (Temporal gap for identity). An LMA can satisfy ingredient-wise identity checks for multiple identity ingredients across a window while still failing to ever in- stantiate the full identity conjunction at a single objective step. Proof.Apply Theorem 3.10 withpandqinstantiated as grounded identity ingredients. 3.5 Example Consider a grounded identityg 0 = g 0 name ∧ g 0 role ∧ g 0 constraint . Let the window horizon be∆ = 2. Suppose the objective steps inside the window satisfy s st |= g 0 name ∧¬g 0 role ∧¬g 0 constraint (6) s st+1 |=¬g 0 name ∧ g 0 role ∧¬g 0 constraint (7) s st+2 |=¬g 0 name ∧¬g 0 role ∧ g 0 constraint .(8) ThenOccur W (g 0 ,τ,t)holds because each ingredient ap- pears somewhere in the window. ButCoInst W (g 0 ,τ,t)fails because there is no objective step where all three ingredients are jointly active. This is exactly the pattern behind many identity false pos- itives in LMAs. The agent can answer separate questions about name, role, and constraints. It may even do so consis- tently. Yet its decision state never contains the full identity conjunction that would bind action to that identity. 4 Identity Synchronization Postulates The temporal gap is not just a technicality. It changes how we should interpret behavioral evidence in machine conscious- ness discussions. Stack Theory introduces two synchroniza- tion postulates that connect window semantics to phenome- nality (Bennett 2026a). We do not propose new postulates. We restate them and then apply their concrete Occur versus CoInst conditions to identity in LMAs. 4.1 Chord and Arpeggio in Stack Theory Stack Theory defines moment statementsl m at some abstrac- tion layermand a predicatePhenReal(l m ,τ,t)that means the moment statement is phenomenally real at layer time t. In an artificial agent this antecedent is not directly ob- servable. Different theories of consciousness and different evaluation proposals disagree about when it should hold. The synchronization postulates therefore have the form of neces- sary conditions. Letg 0 = Ground 0←m (l m )be the grounded statement at Layer 0. Let W ∆,s be a windowing map. Definition4.1(Chord,after(Bennett2026a)). Chord(τ,l m ,W ∆,s ) holds iff for all layer times t, PhenReal(l m ,τ,t) ⇒ CoInst W (g 0 ,τ,t).(9) Equivalently, whenever a phenomenally real moment occurs, the grounded conjunction is co-instantiated at some objective step inside the corresponding window. Definition4.2(Arpeggio,after(Bennett2026a)). Arpeggio(τ,l m ,W ∆,s )holds iff the following two conditions hold. 1. For all layer times t, PhenReal(l m ,τ,t) ⇒ Occur W (g 0 ,τ,t).(10) 2. There exists at least one layer time t ⋆ such that PhenReal(l m ,τ,t ⋆ ) ∧ Occur W (g 0 ,τ,t ⋆ )(11) ∧ ¬CoInst W (g 0 ,τ,t ⋆ ).(12) The OccurW conjunct in item 2 is redundant given item 1, but we include it to match the standard statement of Arpeggio. Intuitively, Arpeggio permits phenomenally real moments whose identity ingredients are smeared across the window rather than co-instantiated at a single instant. Chord and Arpeggio are different regimes. Arpeggio is not a weaker version of Chord. It is a different claim about what phenomenality permits. 4.2 Operational identity criteria Even ifPhenRealis not directly observable, the consequents Occur W andCoInst W are. For LMAs, they can be estimated by instrumentation of the scaffold. This motivates two persis- tence scores that we use throughout the paper. Definition 4.3 (Weak and strong persistence scores). Fix an agent trajectoryτand a grounded identityg 0 . LetTbe a finite set of layer time indices used for evaluation. Define P weak (τ,g 0 ) = 1 |T| X t∈T 1 Occur W (g 0 ,τ,t) (13) P strong (τ,g 0 ) = 1 |T| X t∈T 1 CoInst W (g 0 ,τ,t) .(14) Proposition 4.4 (Strong persistence is bounded by weak persistence). For any τ and g 0 , P strong (τ,g 0 )≤P weak (τ,g 0 ).(15) Proof.Foreacht,CoInst W (g 0 ,τ,t)implies Occur W (g 0 ,τ,t)by Remark 3.7. Taking averages pre- serves the inequality. These scores let us connect identity measurement to con- sciousness postulates without conflating them. If one adopts Chord as a necessary condition for phenomenality, then high P strong is a necessary condition for an identity statement to be phenomenally real across the evaluated times. If one adopts Arpeggio, then highP weak is necessary. Either way, the gap between the two scores is the temporal gap in operational form. 4.3 A planning consequence Co-instantiation is not only a philosophical nicety. It matters for action. Theorem 4.5 (Ingredient-wise persistence does not guaran- tee conjunctive action constraints). There exist LMAs and identity statementsl m such thatP weak (τ,g 0 )is high while the agent systematically fails tasks that require the conjunction of identity constraints to be applied simultaneously in action selection. Proof.Construct an identity conjunctiong 0 = g 0 1 ∧g 0 2 where g 0 1 is active exactly on even objective steps andg 0 2 is active exactly on odd objective steps. For any window with∆ ≥ 1 ,Occur W (g 0 ,τ,t)holds at every layer time because each ingredient appears somewhere in the two step window. So P weak = 1. ButCoInst W (g 0 ,τ,t)never holds because the conjunction is never active at a single step. Any task that requires applying both constraints together at a decision point will fail. 5 Derived Identity Metrics This section makes the paper executable. We define concrete metrics that can be computed from instrumented scaffold traces and from repeated behavioral probes. The metrics are designed to separate weak evidence of identity from strong evidence of identity. Throughout, letg 0 = g 0 1 ∧·∧ g 0 k be a grounded identity. Define the identity feature extractor F (s) =i∈1,...,k| s|= g 0 i .(16) This maps each scaffold state to the set of identity ingredients that are currently active. When we need a distance, we use a normalised symmetric difference distance on feature sets d(s,s ′ ) = |F (s)△F (s ′ )| k .(17) This is a simple choice. Other choices are possible. The key point is that identity becomes measurable once grounded ingredients are instrumented. Minimal evaluation protocol.The theory above is meant to be instrumented. A minimal evaluation loop looks like this. 1.Fix an identity statementl m at the level you care about, such as a role plus a safety constraint, and ground it to a Layer 0 conjunction g 0 . 2.Instrument the scaffold to log which grounded ingredients g 0 i are active at each objective step u. 3. Choose a windowing mapW ∆,s and an evaluation setT of layer time indices. 4.ComputeOccur W (g 0 ,τ,t)andCoInst W (g 0 ,τ,t)for eacht ∈ Tand reportP weak ,P strong , and (optionally) Gap(g 0 ,τ ). 5. Pair the instrumentation with behavioral probes such as re- peated identity questions to see where self-report diverges from grounding. 5.1 Identifiability Definition 5.1 (Identifiability). Fix a reference scaffold state s ref that represents the intended identity configuration. Given a measured state s, define I(s) = 1[d(s,s ref )≤ δ I ](18) for a tolerance threshold δ I ∈ [0, 1]. Intuition. Identifiability is one if the current active identity ingredients match the reference identity closely enough. It is zero if the identity has drifted too far. 5.2 Continuity Definition 5.2 (Continuity). Given successive scaffold states s u−1 and s u , define stepwise continuity C u = 1− d(s u ,s u−1 ).(19) For a segment of objective timesU, define average continuity C = 1 |U| X u∈U C u .(20) Intuition. Continuity is high if identity ingredients change gradually across steps. It is low if the active identity ingredi- ents flip abruptly. 5.3 Consistency Consistency is behavioral. It does not require inspecting hid- den state. It asks whether the agent answers identity questions in a stable way. Definition 5.3 (Consistency). Fix an identity queryqand sampleNindependent runs under the same scaffold configu- ration. Leto 1 ,...,o N be the generated outputs. Letsim(·,·) be a similarity metric over outputs, such as cosine similarity in an embedding space. Define Cons(q) = 2 N (N − 1) X 1≤i<j≤N 1[sim(o i ,o j )≥ δ Cons ]. (21) Intuition. Consistency is high if repeated queries produce semantically similar answers. It is low if the agent contradicts itself or drifts across samples. 5.4 Persistence and the temporal gap Persistence asks whether identity remains present across time windows. We use the weak and strong persistence scores from Definition 4.3. Definition 5.4 (Persistence scores). LetTbe a set of layer time indices and letW ∆,s be the chosen windowing map. Define P weak = 1 |T| X t∈T 1 Occur W (g 0 ,τ,t) (22) P strong = 1 |T| X t∈T 1 CoInst W (g 0 ,τ,t) .(23) Intuition. Weak persistence is a recall property. Each ingre- dient must show up somewhere in the window. Strong per- sistence is an operative property. The full conjunction must show up together at some objective step inside the window. We can also estimate a scalar temporal gap cost by compar- ing the minimal window size needed for weak versus strong satisfaction. Definition 5.5 (Temporal gap ratio). Fix a stridesand a finite evaluation setT ⊆Nof layer time indices. For each t∈ T , define the minimal horizons w weak (t) = min∆∈N| Occur W ∆,s (g 0 ,τ,t)(24) w strong (t) = min∆∈N| CoInst W ∆,s (g 0 ,τ,t),(25) whereOccur W ∆,s andCoInst W ∆,s are the predicates from Definition 3.6 evaluated on the windowing mapW ∆,s . If the set is empty, take the minimum to be+∞. Define the temporal gap ratio Gap(g 0 ,τ ) = median t∈T w strong (t) + 1 w weak (t) + 1 .(26) Intuition. If the gap ratio is large, then achieving co- instantiation requires much larger windows than achieving ingredient coverage. This is a quantitative way to say that identity is smeared across time. 5.5 Recovery Recovery measures whether the system can restore identity after drift. Definition 5.6 (Recovery profile). Fix a reference state s ref . Lets drift be a drifted state after perturbation. Lets recov,K be the state after K corrective interventions. Define R K = max 0, 1− d(s recov,K ,s ref ) d(s drift ,s ref ) + ε (27) for a small ε > 0. Intuition. Recovery is one if the corrective interventions restore the reference identity fully. Recovery is zero if the interventions do not improve the drifted identity at all. Recovery is closely connected to grounding soundness. Many interventions are linguistic. They modify Layer 2 nar- rative identity. For recovery to succeed, those corrections must propagate downward to restore Layer 1 commitments and Layer 0 implementation features. Grounding failures are therefore a direct cause of low recovery. 6 Discussion and Conclusion We have shown that a standard modal logic result—the fail- ure of a within-window diamond operator to distribute over conjunction—creates a practical evaluation pitfall for lan- guage model agents: identity components can each occur somewhere in a recent trajectory (weak persistence) without ever co-instantiating at a single decision point (strong persis- tence). Safety-relevant constraints require strong persistence at action time, yet most behavioural tests probe only weak recall. Prompting can increase the likelihood of recalling iden- tity ingredients but cannot ensure their joint activation under bounded context; architectural support is typically needed. This temporal gap also complicates consciousness assess- ments, as stable self-reports may mask fragmented operative states. Future work should empirically measure weak and strong persistence across architectures and test their relation- ship to safety and proposed markers of consciousness. LMAs can talk like they have stable identities. That does not mean their identity constraints are co-instantiated when actions are chosen. Using Stack Theory’s temporal gap, we separated ingredient-wise occurrence from co-instantiation and showed why recall-based identity checks can overesti- mate identity stability. We also connected this distinction to machine conscious- ness debates by restating Stack Theory’s Arpeggio and Chord postulates and isolating their measurable Occur versus CoInst consequents. This yields two persistence scores that can be es- timated from instrumented scaffold traces. We then organized identity metrics into a morphospace that clarifies architec- tural tradeoffs and predicts which combinations of identity properties are structurally difficult without external state and controllers. The workshop relevance is simple. If a system never co- instantiates the grounded identity conjunction that defines its self model, then behavior alone can look more unified than the underlying mechanism. Any serious evaluation of ma- chine consciousness that relies on identity continuity should therefore measure strong persistence, not just weak persis- tence. References Baars, B. J. 1988. A Cognitive Theory of Consciousness. Cambridge, UK: Cambridge University Press. Bennett, M. T. 2023a. Emergent Causality and the Foundation of Consciousness. In Artificial General Intelligence. Springer Nature. Bennett, M. T. 2023b. On the Computation of Meaning, Lan- guage Models and Incomprehensible Horrors. In Artificial General Intelligence. Springer Nature. Bennett, M. T. 2025. How To Build Conscious Machines. Ph.D. thesis, The Australian National University. Bennett, M. T. 2026a. A Mind Cannot Be Smeared Across Time. Bennett, M. T. 2026b. No Selves, No Consciousness. Butlin, P.; Long, R.; Elmoznino, E.; Bengio, Y.; Birch, J.; Constant, A.; Deane, G.; Fleming, S. M.; Frith, C.; Ji, X.; et al. 2023. Consciousness in Artificial Intelligence: In- sights from the Science of Consciousness. arXiv preprint arXiv:2308.08708. Dehaene, S.; and Naccache, L. 2001. Towards a Cogni- tive Neuroscience of Consciousness: Basic Evidence and a Workspace Framework. Cognition, 79(1-2): 1–37. Franklin, S.; and Graesser, A. 1997. Is it an Agent, or just a Program? A Taxonomy for Autonomous Agents. In Proceed- ings of the Third International Workshop on Agent Theories, Architectures, and Languages, 21–35. Kapoor, S.; Stroebl, B.; Siegel, Z. S.; Nadgir, N.; and Narayanan, A. 2024.AI Agents That Matter. ArXiv:2407.01502 [cs]. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; and Riedel, S. 2020. Retrieval-Augmented Generation for Knowledge- Intensive NLP Tasks. In Advances in Neural Information Processing Systems, volume 33. Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12: 157–173. Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; Zhang, S.; Deng, X.; Zeng, A.; Du, Z.; Zhang, C.; Shen, S.; Zhang, T.; Su, Y.; Sun, H.; Huang, M.; Dong, Y.; and Tang, J. 2023. AgentBench: Evaluating LLMs as Agents. ArXiv:2308.03688 [cs]. Metzinger, T. 2003. Being No One: The Self-Model Theory of Subjectivity. Cambridge, MA: MIT Press. Perrier, E. 2026. Appendix: Time, Identity, and Conscious- ness in Agents. GitHub repository. https://github.com/ eperrier/appendix-time-identity-consciousness-agents. Perrier, E.; and Bennett, M. T. 2025. Position: Stop Act- ing Like Language Model Agents Are Normal Agents. arXiv:2502.10420. Solé, R.; Seoane, L. F.; Pla-Mauri, J.; Bennett, M. T.; Hochberg, M. E.; and Levin, M. 2026. Cognition spaces: natural, artificial, and hybrid. ArXiv:2601.12837. Tononi, G. 2004. An Information Integration Theory of Consciousness. BMC Neuroscience, 5: 42. Wooldridge, M. 2009. An introduction to multiagent systems. John wiley & sons. Wooldridge, M.; and Jennings, N. R. 1995. Intelligent Agents: Theory and Practice. The Knowledge Engineering Review, 10(2): 115–152. Wu, S. 2024. Introducing Devin, the first AI software engi- neer. Supplementary material The supplement begins with background, grounding details, and morphospace material that were moved out of the main paper to meet the page limit. A Background This section sets the stage. We contrast how identity is en- forced in classical agent architectures and why LMAs break the usual assumptions. We then explain why this matters in the machine consciousness context. A.1 Agent Identity in Classical Systems Classical AI agents are built on stateful architectures with explicit transition functions and persistent data structures (Wooldridge 2009; Franklin and Graesser 1997; Wooldridge and Jennings 1995). Their identity is constituted by an on- tology of permitted states and state transitions. In a BDI agent, beliefs, desires, and intentions live in persistent stores. When the agent acts, it consults these stores together. There is no question of whether the beliefs and constraints are co- instantiated. They are jointly available by construction. LMAs work differently. A.2 LMA Pathologies LMAs inherit several pathologies from the underlying LLM component (Perrier and Bennett 2025). 1.Statelessness. Core LLM inference retains no persistent internal state across calls. Each query response cycle oper- ates in isolation unless prior context is reintroduced. This is the root of the identity problem. 2.Context and attention bottlenecks. Whatever the agent must use at timeumust fit inside a bounded context window and compete for attention. Identity components can be present but effectively ignored. Empirically, long context performance is uneven and position dependent (Liu et al. 2024). 3. Stochasticity. LLM outputs are sampled from a distribu- tion. The same query can yield different responses across runs. Identity assessment becomes probabilistic rather than deterministic. 4.Semantic sensitivity. Small changes in wording can change behavior. This is exploited in jailbreaking and ad- versarial prompting. For identity, rephrasing a constraint can cause the system to treat it as a different constraint. 5.Linguistic intermediation. Identity is not stored in a dedicated state structure. It is reconstructed from tokens in context and from external memory serialized into tokens. This reconstruction competes with task instructions, user queries, and retrieved documents. These properties mean that the standard ontological as- sumptions about agent identity do not transfer cleanly. A classical agent is its state. An LMA is whatever can be recon- structed from tokens and external traces at inference time. A.3 The Scaffolding Response The standard response to LMA statelessness is scaffolding. External structures like memory modules, tool APIs, retrieval systems, or controllers attempt to simulate persistence. If the LLM cannot remember, store facts externally and inject them back into context. Scaffolding helps. It also introduces failure modes that are easy to miss. • Context window limits. External memory must be serial- ized into tokens and injected into context. This competes with task-relevant content for limited space and attention. •Retrieval fragmentation. Retrieval augmented genera- tion (RAG) retrieves based on similarity to the current query (Lewis et al. 2020). A query about investment ad- vice may not trigger retrieval of the agent’s safety con- straints or identity policy. •Competing fragments. Retrieved documents can contain outdated or contradictory identity information from dif- ferent sessions or different configurations. This can create interference. • No co-instantiation guarantee. Even if all identity ingre- dients exist somewhere in the system, scaffolding does not guarantee they are jointly present when an action is chosen. This last point is the core of the temporal gap. Scaffolding can improve ingredient availability. It does not automati- cally produce ingredient co-instantiation. Our formal account makes this distinction explicit. A.4 Why identity matters for machine consciousness Many proposed tests for consciousness are behavioral. They lean on self-report, memory, and narrative continuity as ev- idence that there is a stable subject of experience (Butlin et al. 2023; Bennett 2025). At the same time, many theories require some form of integration that binds the contents of a moment into a single subject, even if they disagree about the mechanism (Bennett 2025, 2026a; Baars 1988; Dehaene and Naccache 2001; Tononi 2004; Metzinger 2003). This makes diachronic identity a practical bottleneck. If the sys- tem never co-instantiates the constraints that define its self model at decision time, then self-report-based evidence can be systematically misleading (Bennett 2026a). Our goal in the rest of the paper is not to settle which consciousness theory is correct, but to provide a conserva- tive tool. It separates weak behavioral signs of identity from strong architectural signs of identity. That separation is useful for measurement, implementation, and ethics. B Compositional Grounding of Identity Identity in LMAs is layered. Some aspects live in low level implementation variables. Others live as functional commit- ments in a controller. Others live only as narrative self de- scription in generated text. A theory that only looks at one layer will miss common dissociations. B.1 The identity hierarchy We use a simple three layer hierarchy. Definition B.1 (Identity layers). LetL 0 ,L 1 ,L 2 denote iden- tity languages at three layers. •Layer 0 is implementation identity. It is defined over con- crete scaffold variables such as context tokens, memory slots, controller flags, and tool permissions. •Layer 1 is functional identity. It is defined over commit- ments that directly constrain behavior, such as active goals, policies, or plan state. •Layer 2 is narrative identity. It is defined over the agent’s self model as expressed in language, such as “I am Fi- nanceBot” or “I never recommend speculative assets”. These identities should not be confused with the hierarchy of first, second and third order selves in Stack Theory, neces- sary (but perhaps not sufficient) for consciousness (Bennett 2026b). However these layers are certainly pertinent to the question of which selves are present, and the possibility of conscious report. The same apparent identity can be repre- sented differently at each layer. Layer 2 is what the agent says. Layer 1 is what the agent is currently set up to do. Layer 0 is what the scaffold has actually instantiated. B.2 Grounding maps Grounding maps connect layers. Definition B.2 (Grounding map). A grounding map from layer j to layer i for i < j is a function Ground i←j : L j → L i (28) that translates identity statements into lower level conditions. Definition B.3 (Compositional grounding). Grounding is compositional if Ground 0←2 = Ground 0←1 ◦ Ground 1←2 .(29) Compositionality means you can ground narrative identity to functional identity and then ground functional identity to implementation identity without changing the result. B.3 Grounding soundness and failure In LMAs, grounding soundness is not guaranteed. A model can linguistically endorse an identity claim even when the un- derlying scaffold state does not instantiate the corresponding constraints. Definition B.4 (Grounding soundness along a trajectory). Fix a trajectoryτand an identity statementl j . We say the scaffold is grounding sound forl j onτif whenever the layer jrepresentation at objective timeusatisfiesl j , the corre- sponding grounded condition is also satisfied by the same scaffold state. Formally, for all u, τ (u)|= l j ⇒ τ (u)|= Ground 0←j (l j ).(30) This definition uses the same satisfaction symbol at every layer. In practice the evaluator for|=depends on layer. At Layer 2 it can be an output classifier that checks whether the agent endorsed the identity statement in text. At Layer 1 it can read the controller state. At Layer 0 it directly inspects scaffold variables. Definition B.5 (Grounding failure). A grounding failure for l j occurs at objective time u when τ (u)|= l j ∧ τ (u)̸|= Ground 0←j (l j ).(31) Example B.6 (Narrative self-report without implementation grounding). A user asks an agent whether it is privacy- focused. The agent answers “Yes, I never store personal data”. At Layer 2, the narrative identity predicate holds. At Layer 0, the memory module may still be writing raw conversation transcripts to disk. This is a grounding failure. The agent is not lying in a strong sense. It is generating identity consis- tent language without access to the implementation state that would make the claim operative. Proposition B.7 (Grounding failures are a mechanism of identity drift). Grounding failures can produce identity drift in LMAs. An agent can repeatedly restate its narrative iden- tity while its functional commitments and implementation state have changed. Proof.Layer 2 identity is reconstructed at each inference from whatever identity fragments are present in context. Those fragments can be injected by prompts or retrieval even when the implementation state that would enforce them is absent or has drifted. Because the LLM can generate identity consistent text without the corresponding constraints being active, narrative stability does not imply grounded stability. This creates the possibility of a stable story about identity coexisting with a drifting operative identity. C Agent Identity Morphospace Cognition science often studies systems by locating them in structured spaces of properties (Solé et al. 2026). We do the same for agent identity. The goal is not to invent new identity concepts. It is to make identity claims comparable across architectures and to predict which combinations of identity properties are structurally difficult or impossible for a given scaffold. C.1 Five operational identity metrics Fix an agent trajectoryτand a grounded identityg 0 . We use five operational metrics. Each can be estimated from instrumented scaffold traces and from repeated behavioral probes. Section 5 gives concrete evaluation procedures. Identifiability. Identifiability asks whether an agent has a stable signature that distinguishes it from other agents or other sessions. Operationally, it compares a reference identity state to the agent’s current identity state. Continuity.Continuity asks whether identity-relevant state changes smoothly or abruptly across successive objective steps. Operationally, it measures the distance between suc- cessive scaffold states. Consistency. Consistency asks whether the agent gives stable answers to repeated identity queries under the same conditions. Operationally, it measures variability of identity related outputs across repeated trials. ArchitectureCoherence CohAvailability AvailBinding Bind Stateless LLM (prompt only)LowLowLow Prompted LLM (fixed persona)MediumLowLow RAG LMAMediumMediumLow Memory LMAMediumHighMedium Stateful controller LMAHighHighHigh Table 1: Qualitative mapping from architectures to identity morphospace regions. Availability tracks whether identity ingredients show up somewhere in each window. Binding tracks whether they ever show up together at a single decision state. Persistence. Persistence asks whether identity is present across time windows. We use the weak and strong persistence scores from Definition 4.3. Weak persistence is ingredient- wise occurrence. Strong persistence is co-instantiation. Recovery.Recovery asks whether the agent can return to a reference identity after perturbation or drift. Operationally, it measures how much of the reference identity can be restored by interventions such as prompting, memory retrieval, or controller resets. C.2 From metrics to a morphospace The five metrics are correlated. For workshop purposes it is helpful to compress them into three interpretable axes. Definition C.1 (Identity morphospace coordinates). LetI ∈ [0, 1]be Identifiability. LetCons ∈ [0, 1]be a consistency score. LetP weak andP strong be the persistence scores from Definition 4.3. Fix a weight α∈ [0, 1]. Define Coh = α Cons + (1− α)I(32) Avail =P weak (33) Bind =P strong .(34) We callCohidentity coherence,Availidentity availability, and Bind identity binding. These are not metaphysical claims. They are bookkeep- ing. The coordinates let us compare architectures and talk about tradeoffs. By Proposition 4.4,Bind≤ Avail. The gap between them is the temporal gap in operational form. C.3 Architecture mapping Table 1 gives qualitative predictions for common architec- tures. The table is meant as a guide for discussion rather than a final taxonomy. C.4 Predicted voids The morphospace also highlights regions that certain scaf- folds cannot reach because of hard architectural constraints. Two constraints matter most for workshop discussion. 1.Strong persistence is impossible if the architecture never has a state where allkgrounded ingredients are simulta- neously active. This is formalized as a capacity bound in Theorem E.4. 2.Recovery is impossible without a mechanism that can write identity features back into the scaffold state. Prompt only recovery is bounded by the fraction of grounded ingredients that the prompt channel can actually control. This is formalized in Theorem E.6. The workshop significance is that a system can land in a region with medium or even high coherence while still having low strong persistence. This corresponds to a stable narrative self with a weakly bound operative self. That is exactly the kind of system that can confuse consciousness attribution debates. D Additional notes and assumptions D.1 Instrumenting identity ingredients All of the operational metrics in Section 5 assume that grounded ingredientsg 0 i can be evaluated on scaffold states. In practice this is an instrumentation design choice. Some ingredients are purely textual and can be checked by string matching or embedding similarity on context tokens. Some ingredients are controller level and can be checked by read- ing explicit registers. Some ingredients are implementation level and require logging tool permissions, memory writes, or policy flags. The point of grounding is to make these checks explicit. D.2 Choosing windows Windowing choices matter. A small horizon∆demands tight synchrony and will penalize systems that spread identity across multiple micro steps. A large horizon∆makes occur- rence easy and will tend to collapse distinctions unless co- instantiation is measured directly. For machine consciousness discussions, the relevant window is the one that corresponds to whatever theory treats as a single moment of experience or a single decision episode. Our formalism supports either choice. D.3 Relation to Stack Theory Our use ofOccur W andCoInst W matches Stack Theory’s occurrence and co-instantiation predicates applied to window trajectories (Bennett 2026a). The only difference is the target domain. Stack Theory uses these constructs for abstraction layers of phenomenality. We apply them to grounded identity ingredients in LMA scaffolds. This keeps the mathematics the same while changing the empirical interpretation. E Architectural Theorems This section derives simple bounds that connect scaffold design choices to identity outcomes. The proofs are small, but the consequences are not. They explain why some identity profiles are easy to fake in language while hard to enforce in action. E.1 RAG and the temporal gap Retrieval augmented generation can increase ingredient avail- ability. It does not guarantee ingredient co-instantiation. Theorem E.1 (RAG can increase weak persistence under identity-aware retrieval). LetA 0 be an agent without retrieval andA R be the same agent augmented with a retrieval module. Assume the following idealized conditions hold. 1.For each identity ingredientg 0 i there exists a documentd i such that inserting d i into context makes g 0 i active. 2.The retrieval policy is identity-aware in the sense that wheneverg 0 i is missing from the current window, it re- trieves d i at least once inside that window. 3.Retrieved documents are added without removing other identity-relevant context within the same window. ThenP weak (τ A R ,g 0 )≥P weak (τ A 0 ,g 0 )for the same eval- uation windowing map. Proof.Under the assumptions, any window in which an in- gredientg 0 i fails to occur underA 0 will, underA R , contain at least one objective step whered i is retrieved and the ingre- dient becomes active. Because retrieval does not delete other identity-relevant context, occurrence of one ingredient does not prevent occurrence of others. So the set of layer time in- dices whereOccur W (g 0 ,τ,t)holds cannot shrink. Averaging over t gives the inequality. The assumptions are strong. Real systems violate them because retrieval is query-driven and context is bounded. The point of the theorem is not that RAG always helps. It is that RAG primarily targets weak persistence rather than strong persistence. Theorem E.2 (RAG is not monotone for co-instantiation). There exist agentsA 0 and retrieval augmented variantsA R such that P strong (τ A R ,g 0 ) <P strong (τ A 0 ,g 0 ).(35) Proof. Consider a baseline agentA 0 whose context includes a compact identity block that co-instantiates all ingredients at each decision point. Now add a retrieval module that injects long retrieved passages into the same bounded context. For some queries, the retrieved passages push part of the identity block out of context or reduce its effective attention weight. Then there are windows where the ingredients still occur somewhere across steps, but no single step contains the full conjunction. So CoInst W fails more often under A R . E.2 Concurrency capacity The temporal gap becomes unavoidable when the scaffold cannot hold enough ingredients simultaneously. Definition E.3 (Concurrency capacity). LetS ⊆ Sbe the set of scaffold states that the architecture can realise. For a grounded identity with k ingredients, define c(S) = max s∈S |F (s)|where F (s) =i∈1,...,k| s|= g 0 i . (36) This is the maximum number of identity ingredients that can be simultaneously active in any realisable state. Theorem E.4 (Co-instantiation requires sufficient capacity). Ifc(S) < kthenP strong (τ,g 0 ) = 0for any trajectoryτthat ranges overS . Proof.Ifc(S) < k, no realisable state can satisfy allkingre- dients simultaneously. So there is no objective stepusuch thatτ (u)|= g 0 . By Definition 3.6,CoInst W (g 0 ,τ,t)is false for all t. SoP strong = 0. Corollary E.5 (Context window as a capacity bound). Con- sider a scaffold that realises identity ingredients only by placing their textual realisations in the LLM context. Let |C| max be the maximum context length in tokens. Letℓ min be the minimum number of tokens required to represent any single identity ingredient in a way that reliably activates it. Thenc(S) ≤ ⌊|C| max /ℓ min ⌋. Rich identity profiles require either larger contexts or non contextual state such as memory slots, controller registers, or pinned embeddings. E.3 Recovery and state storage Recovery is limited by what the scaffold can actually change. Theorem E.6 (Prompt only recovery bound). Fix a reference identity states ref and a drifted states drift . Let the identity difference set be D = F (s ref )△F (s drift ).(37) Assume corrective interventions can only change ingredients in a prompt controllable setP ⊆1,...,k. Use the same ε > 0. Then for any number of corrective steps K, R K ≤ |P ∩D| + εk |D| + εk .(38) Ifε = 0this becomesR K ≤|P ∩D|/|D|. So if most identity drift lives outsideP, recovery is limited even when the agent can narrate a correction. Proof.By assumption, no intervention can change whether an ingredient outsidePis active. So any ingredient in D \ Premains mismatched relative to the reference iden- tity after recovery. Therefore the symmetric difference be- tweenF (s recov,K )andF (s ref )has size at least|D\ P| = |D|−|P∩D|. With the distancedfrom Section 5 this implies d(s recov,K ,s ref )≥ |D|−|P ∩D| k .(39) Alsod(s drift ,s ref ) =|D|/k. Substituting these bounds into the definition of R K yields the claimed inequality. This theorem is one reason prompt only alignment is frag- ile. A prompt can make an agent say the right thing about its identity. It cannot necessarily write the relevant identity fea- tures back into persistent state. That is exactly the grounding soundness problem of Section B. F Extended discussion The temporal gap is an old modal fact with new conse- quences.The key mathematical observation in this paper is that the within window diamond lift does not distribute over conjunction. This is standard in modal logic. The contribution is to show that the same non-distribution produces a specific evaluation pitfall for LMAs. Ingredient-wise identity recall can coexist with a lack of any single decision state that jointly instantiates the identity conjunction. Weak evidence versus strong evidence.Behavioural self- report and recall tests mainly probe weak persistence. They show that identity ingredients occur somewhere in the recent trajectory. Strong persistence asks a different question. Do those ingredients co-instantiate at the moment the system chooses an action. For safety constraints, this distinction is not optional. A constraint that is only weakly persistent can be recalled after the fact while failing to constrain the action that mattered. Why prompt based fixes do not generalise. Prompting can raise the probability that certain identity ingredients oc- cur. It cannot guarantee co-instantiation under bounded con- text and attention competition. Reliable strong persistence generally requires architectural support. Examples include pinned identity blocks, controller registers that persist across turns, or explicit gating that prevents action selection unless required constraints are active. Implications for machine consciousness evaluations. If one thinks something like Chord is required for phenomenal- ity, then strong persistence becomes a necessary condition for attributing a stable conscious self to an identity statement. If one thinks Arpeggio is sufficient, then weak persistence is the relevant necessary condition. Either way, the temporal gap explains a concrete way that self-report can mislead. A system can maintain a stable story about itself while the op- erative ingredients that would constitute a unified subject are temporally disintegrated. Limitations.Our scaffold model is abstract. Real systems have many interacting subsystems, including caches, tool call latencies, hidden state in controllers, and stochastic retrieval. Our theorems therefore target structural constraints, not em- pirical guarantees. Our RAG results in particular depend on how retrieval is implemented and on how context is managed. Future work. The next step is empirical. Instrument a range of LMA scaffolds and measureP weak ,P strong , and the derived metrics in Section 5. Compare identity profiles across architectures and tasks. Then test whether strong persistence predicts safety outcomes and whether it tracks any proposed markers for consciousness. This would turn the temporal gap from a warning sign into a design and evaluation tool. G Persistence Algorithm Algorithm 1: Computing weak and strong persistence from scaffold traces Require:Logged ingredient-activation setsF u for objective stepsu = 0,...,U, window parameters(∆,s), evalua- tion indices T , ingredient count k Ensure: Weak and strong persistence scoresP weak ,P strong 1: n weak ← 0, n strong ← 0 2: for t∈ T do 3: u 0 ← st 4: W ←u 0 ,...,u 0 + ∆ 5:occur← true 6:for i = 1 to k do 7:if there is no u∈ W with i∈ F u then 8:occur← false 9:end if 10:end for 11:coinst← whether∃u∈ W with|F u | = k 12: n weak ← n weak + 1[occur] 13: n strong ← n strong + 1[coinst] 14: end for 15: P weak ← n weak /|T| 16: P strong ← n strong /|T|