Paper deep dive
Context Cartography: Toward Structured Governance of Contextual Space in Large Language Model Systems
Zihua Wu, Georg Gartner
Intelligence
Status: succeeded | Model: anthropic/claude-sonnet-4.6 | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/24/2026, 3:36:31 AM
Summary
This paper introduces Context Cartography, a formal framework for governing contextual space in Large Language Model (LLM) systems. The authors define a tripartite zonal model partitioning information into black fog (unobserved), gray fog (stored memory), and visible field (active reasoning surface), and formalize seven cartographic operators (reconnaissance, selection, simplification, aggregation, projection, displacement, and layering) as transformations governing information transitions between zones. The framework is grounded in the salience geometry of transformer attention, addressing phenomena like the 'lost in the middle' effect, entropy accumulation, and attention dilution. Analysis of four contemporary systems (Claude Code, Letta, MemOS, OpenViking) provides evidence that these operators are converging independently across the industry. The paper derives testable predictions and proposes a diagnostic benchmark for empirical validation.
Entities (46)
Relation Signals (39)
Zihua Wu â affiliatedwith â NVIDIA
confidence 99% ¡ Zihua Wu NVIDIA zihuaw@nvidia.com
Georg Gartner â affiliatedwith â TU Wien
confidence 99% ¡ Georg Gartner Research Division Cartography, TU Wien
Georg Gartner â authored â Context Cartography
confidence 99% ¡ We introduce Context Cartography, a formal framework for the deliberate governance of contextual space.
Zihua Wu â authored â Context Cartography
confidence 99% ¡ We introduce Context Cartography, a formal framework for the deliberate governance of contextual space.
Context Cartography â defines â Tripartite Zonal Model
confidence 99% ¡ We define a tripartite zonal model partitioning the informational universe into black fog, gray fog, and the visible field
Black Fog â hasfailuremode â Hallucination
confidence 99% ¡ Black Fog (âŹ): Hallucination. When an agent generates assertions conditioned on uâ⏠as if uâđą, the result is hallucination
Tripartite Zonal Model â partitionsinto â Gray Fog
confidence 99% ¡ partitioning the informational universe into black fog (unobserved), gray fog (stored memory), and the visible field
Tripartite Zonal Model â partitionsinto â Visible Field
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The prevailing approach to improving large language model (LLM) reasoning has centered on expanding context windows, implicitly assuming that more tokens yield better performance. However, empirical evidence - including the "lost in the middle" effect and long-distance relational degradation - demonstrates that contextual space exhibits structural gradients, salience asymmetries, and entropy accumulation under transformer architectures. We introduce Context Cartography, a formal framework for the deliberate governance of contextual space. We define a tripartite zonal model partitioning the informational universe into black fog (unobserved), gray fog (stored memory), and the visible field (active reasoning surface), and formalize seven cartographic operators - reconnaissance, selection, simplification, aggregation, projection, displacement, and layering - as transformations governing information transitions between and within zones. The operators are derived from a systematic coverage analysis of all non-trivial zone transformations and are organized by transformation type (what the operator does) and zone scope (where it applies). We ground the framework in the salience geometry of transformer attention, characterizing cartographic operators as necessary compensations for linear prefix memory, append-only state, and entropy accumulation under expanding context. An analysis of four contemporary systems (Claude Code, Letta, MemOS, and OpenViking) provides interpretive evidence that these operators are converging independently across the industry. We derive testable predictions from the framework - including operator-specific ablation hypotheses - and propose a diagnostic benchmark for empirical validation.
Tags
Links
- Source: https://arxiv.org/abs/2603.20578v1
- Canonical: https://arxiv.org/abs/2603.20578v1
Trouble viewing inline? Open PDF directly â
Full Text
106,682 characters extracted from source content.
Expand or collapse full text
Context Cartography: Toward Structured Governance of Contextual Space in Large Language Model Systems Zihua Wu NVIDIA zihuaw@nvidia.com Georg Gartner Research Division Cartography, TU Wien georg.gartner@tuwien.ac.at (March 2026) Abstract The prevailing approach to improving large language model (LLM) reasoning has centered on expanding context windows, implicitly assuming that more tokens yield better performance. However, empirical evidenceâincluding the âlost in the middleâ effect and long-distance relational degradationâdemonstrates that contextual space exhibits structural gradients, salience asymmetries, and entropy accumulation under transformer architectures. We introduce Context Cartography, a formal framework for the deliberate governance of contextual space. We define a tripartite zonal model partitioning the informational universe into black fog (unobserved), gray fog (stored memory), and the visible field (active reasoning surface), and formalize seven cartographic operatorsâreconnaissance, selection, simplification, aggregation, projection, displacement, and layeringâas transformations governing information transitions between and within zones. The operators are derived from a systematic coverage analysis of all non-trivial zone transformations and are organized by transformation type (what the operator does) and zone scope (where it applies). We ground the framework in the salience geometry of transformer attention, characterizing cartographic operators as necessary compensations for linear prefix memory, append-only state, and entropy accumulation under expanding context. An analysis of four contemporary systems (Claude Code, Letta, MemOS, and OpenViking) provides interpretive evidence that these operators are converging independently across the industry. We derive testable predictions from the frameworkâincluding operator-specific ablation hypothesesâand propose a diagnostic benchmark for empirical validation. Keywords: context engineering, large language models, agent architecture, memory systems, cartographic generalization, attention mechanisms 1 Introduction The dominant approach to improving LLM reasoning over the past several years has been to expand context windowsâfrom a few thousand tokens to over one million (Ding et al., 2024). This expansion widens the modelâs field of view but does not govern how information moves between sensing, memory, and reasoning; without such governance, visibility becomes overload. Models exhibit severe performance degradation when critical information occupies intermediate positions (Liu et al., 2024), relational knowledge degrades as token distance increases (Li and others, 2025), and effective context utilization falls dramatically short of stated capacity (Paulsen, 2025). The emerging discipline of Context Engineering (Mei et al., 2025) addresses this challenge broadly, encompassing context retrieval, processing, and management. However, it lacks a formal spatial model of how contextual space is structured and why certain management strategies are architecturally necessary. Cartography, understood broadly, is the study of the structured representation of large information spaces under bounded cognition (MacEachren, 1995; Robinson and Petchenik, 1976). While historically grounded in geographic map-making, cartographic principlesâselection, generalization, controlled distortion, scale-dependent representationâapply wherever a bounded medium must faithfully represent an unbounded reality (Gartner et al., 2007): cache hierarchies decide which objects to retain under capacity limits, and information visualization projects high-dimensional data onto two-dimensional displays under perceptual constraints. In each case, a bounded medium must represent an unbounded source through controlled selection, compression, and distortion. We introduce Context Cartographyâa framework that extends cartographic generalization theory (McMaster and Shea, 1992) to the governance of contextual space in LLM systems, treating LLM context as the latest instance of this general bounded-representation problemânot a transcript to be extended but terrain to be mapped, with boundaries, gradients, and transitions that require deliberate governance. The analogy to the âfog of warââa term originating in Clausewitzâs theory of military uncertainty (von Clausewitz, 1832) and later operationalized in real-time strategy gamesâis instructive. In such games, uncertainty is not binary but stratified: black fog (unexplored territory), gray fog (previously revealed but no longer visible), and the visible field (currently illuminated). Skilled players manage transitions between these zones rather than attempting to illuminate everything simultaneously. LLM-based agents face the same structure of uncertainty across context. Our central claim is that context is not a passive container but a structured spatial field whose governance is architecturally necessary. The specific salience geometry we characterize is grounded in transformer attention, but the zonal model and operators are architecture-general. Sequential architectures exhibit positional sensitivity (transformers via attention non-uniformity (Liu et al., 2024); state space models and linear attention architectures via recency-dominant state compression (Gu and Dao, 2024; Peng and others, 2025)); non-sequential architectures such as diffusion language models (Nie et al., 2025) face different constraints (computational cost scaling with sequence length) that shift operator priorities without eliminating the need for governance. The zonal structure itself persists regardless of architecture: the world exceeds any modelâs capacity, memory requires structure, and reasoning requires a bounded surface. The contributions of this paper are fourfold: 1. We formalize a tripartite zonal model of contextual space with explicit definitions, state transitions, and failure modes (SectionË3.1). 2. We characterize the salience geometry of transformer context windows, establishing why cartographic governance is architecturally necessary (SectionË3.2). 3. We define seven cartographic operatorsâderived from a systematic coverage analysis of all zone transformationsâas formal transformations organized by transformation type and zone scope, with specified cartographic correspondences (SectionË3.3). 4. We derive testable predictions from the framework, propose a diagnostic benchmark, and outline a research agenda for empirical validation (SectionË5). 2 Related Work Long-Context Language Models. Techniques including sparse attention (Beltagy et al., 2020), hardware-aware attention implementation (Dao et al., 2022), rotary positional encodings, and KV cache optimization (Li et al., 2025b) have enabled context lengths exceeding one million tokens (Ding et al., 2024). However, Liu et al. (2024) demonstrated that performance degrades sharply when relevant information occupies intermediate positions. Subsequent work has established that this positional bias emerges from the structural properties of causal attention (Wu et al., 2025). Empirical benchmarks confirm the effect across model families (Gupte et al., 2025). Mitigations such as multi-scale positional encoding (Zhang et al., 2024), hidden-state scaling (Yu et al., 2025), and attention calibration (Hsieh et al., 2024) offer partial remedies but do not address the fundamental geometry of contextual space. Similarly, hierarchical and sparse attention patterns (e.g., Longformer (Beltagy et al., 2020)) and recurrent memory architectures (e.g., RWKV (Peng and others, 2025)) reshape the salience function but do not govern how information moves between zones; Context Cartography operates at the governance layer above architectural attention patterns. Prompt Engineering and Chain-of-Thought. Chain-of-Thought prompting (Wei et al., 2022) demonstrated that structuring the reasoning trace within the context window significantly improves performance on complex tasks. This can be understood as an implicit application of the layering operator: separating the reasoning trace from the task specification within the same sequence. More broadly, prompt engineering operates at the level of local phrasing, while Context Cartography operates at the level of spatial organization across the entire context. Retrieval-Augmented Generation. RAG (Lewis et al., 2020) brings information from the black fog into the visible field through retrieval. Surveys by Gao et al. (2024) catalog the evolution of retrieval strategies. However, retrieval from a pre-indexed corpus operates primarily as a selection and projection mechanismâdetermining what to recall from stored memory without systematically addressing how recalled content should be positioned or layered within the context window (Cuconasu et al., 2024; Shi et al., 2023). Memory-Augmented Agent Architectures. Tool-augmented agents (Schick et al., 2023; Yao et al., 2023) extend LLM capabilities through external function calls. MemGPT (Packer et al., 2023) models context management as an operating system problem. MemOS (Li et al., 2025d) introduces a memory operating system with structured, typed memory containers (MemCubes). Generative Agents (Park et al., 2023) maintain persistent memory streams that agents retrieve and reflect upon. Memorizing Transformers (Wu et al., 2022) augment attention with retrieval over external memory stores. Recent surveys (Liu and others, 2025) identify three dominant memory realizations (token-level, parametric, latent) with distinct functional roles. These systems implicitly adopt principles we formalize as cartographic operators. Context Engineering. Mei et al. (2025) establish Context Engineering as a broad discipline encompassing retrieval, processing, and management of context for LLMs. Context Cartography is complementary: where Context Engineering provides a comprehensive taxonomy of what techniques exist, Context Cartography provides a formal spatial model of why certain techniques are necessary and how they relate to the geometry of contextual space. Specifically, Context Cartography adds: (i) a zonal model partitioning the informational universe into epistemic states, (i) a geometric analysis grounding management strategies in attention physics, and (i) an operator formalism derived from cartographic generalization theoryâenabling distinctions (e.g., archival vs. destructive summarization) that a flat taxonomy does not capture. Cartographic Generalization and Ubiquitous Cartography. Cartographic generalization reduces map information content while preserving essential characteristics (McMaster and Shea, 1992). Standard operators include selection, simplification, aggregation, typification, displacement, collapse, and symbolization (Roth et al., 2011; Stoter et al., 2014). Brassel and Weibel (1988) established the canonical process model for automated generalizationâstructure recognition, process modeling, executionâthat our cartographic pipeline concept (SectionË3.3) extends to contextual space. Robinson and Petchenik (1976) formalized cartographic communication as an encodingâtransmissionâdecoding channel, and MacEachren (1995) integrated cognitive and semiotic approaches to show that maps function as abstract synthetic representations under bounded cognition. Gartner et al. (2007) argue that cartographic principles extend beyond geographic maps to any context where spatial information is communicated under constraintsâa position we adopt and extend to LLM context management. Recent work has noted that spatial hierarchies from cartography improve LLM performance on entity disambiguation tasks (Gartner, 2025), and the adaptive cartographic methodologies surveyed by Lin et al. (2011) for virtual geographic environments anticipate patterns now emerging in agent architecture design. 3 The Context Cartography Framework 3.1 Contextual Space and Epistemic Zones Definition 1 (Contextual Universe). Let U denote the contextual universeâthe totality of information potentially relevant to an agentâs task. At any time step t, U is partitioned into three disjoint epistemic zones: =âŹtâŞtâŞtU=B_t\;âŞ\;G_t\;âŞ\;V_t where âŹtB_t (black fog) is the unobserved frontier, tG_t (gray fog) is stored memory, and tV_t (visible field) is the active reasoning surface. Definition 2 (Zone Transitions). Information moves between zones via four transition functions: sense :âŹâ :B (tool calls, search, code execution) (1) recall :â :G (retrieval, memory loading) (2) evict :â :V (summarization, archival) (3) expire :â⏠:G (staleness, invalidation) (4) The composition recallâsense recall sense maps directly from unobserved territory to the reasoning surface. In minimal agent implementations, sense may deposit results directly into V (bypassing G entirely). We treat direct âŹâB as a collapsed composition of sense followed by immediate recall; Context Cartography asserts that this composition should be mediatedâraw sensing outputs should be routed through G or at minimum transformed before entering V, to prevent visible field contamination. Mediation is most critical when the sensing output is large relative to |||V| or when its format does not match the current projection schema p; for small, pre-structured outputs (e.g., a single API response returning typed fields), the collapsed composition is acceptable. FigureË1 illustrates the tripartite zonal architecture and its transitions. âŹBBlack Fog(Unobserved)GGray Fog(Stored Memory)VVisible Field(Active Context)senserecallevictexpirerecallâ\, \,senseHallucinationDrift, bloatOverload, dilution Figure 1: The tripartite zonal architecture. Solid arrows show the four named transitions; the dashed arc shows the degenerate âŹâB path that Context Cartography recommends mediating through G or operator guards (ĎĎ, Ď+Ď^+). Zone-level failure modes appear below each zone. Each zone is characterized by a distinct failure mode when governance is absent: Black Fog (âŹB): Hallucination. When an agent generates assertions conditioned on uââŹu as if uâu , the result is hallucination (Huang et al., 2023). Although Kadavath et al. (2022) show that models possess partial self-knowledge about their uncertainty, this calibration is unreliable at epistemic boundariesâmodels frequently generate confident continuations consistent with pre-training statistics for information they have not actually observed. Gray Fog (G): Drift and Bloat. Memory is inherently distinct from active visibility: the underlying reality may have changed since last observation (Liu and others, 2025). If G is maintained as a chronological log, retrieval yields unstructured text that overwhelms V upon re-entry (Packer et al., 2023). Zhang et al. (2025b) identify two specific failure modes: brevity bias (optimization toward short, generic outputs) and context collapse (monolithic rewriting eroding accumulated knowledge). We recharacterize both as transition-level failures below. Visible Field (V): Overload and Dilution. The visible field is analogous to the episodic buffer in Baddeleyâs working memory model (Baddeley, 2000)âa limited-capacity interface between perception and long-term memory. Cognitive load theory (Sweller, 1988) predicts that exceeding this capacity degrades performance. In LLMs, the corresponding failure modes are attention dilution (Liu et al., 2024) and constraint drift (system instructions losing influence as context grows (Gu, 2026)). Transition-Level Failure Modes. The zone-level failure modes above arise from the state of each zone; a complementary class of failures arises from how information moves between zones: ⢠Visible field contamination. When the composition recallâsense recall sense is executed without mediationâraw tool output or search results entering V without simplification (ĎĎ) or format transformation (Ď+Ď^+)âthe visible field absorbs noise at the expense of signal. This is not a missing operator but a failure to compose existing operators on the inbound path. ⢠Context collapse via destructive compaction. If ĎâĎ^- is applied without first archiving the original content in G, the pre-compaction context transitions directly to âŹB (permanently lost) rather than G (retrievable). The context collapse and brevity bias identified by Zhang et al. (2025b) are precisely this failure: ĎĎ or ĎâĎ^- applied destructively, bypassing archival. ⢠Cross-agent projection incoherence. In multi-agent settings, each agent maintains its own zonal structure. When agents exchange information, they implicitly apply Ď+Ď^+ across agent boundaries. If their projection schemas are incompatibleâone agentâs summary omits distinctions another agent requiresâthe receiving agentâs V contains well-formed but semantically distorted content. This is supported by systematic failure analyses: Cemri et al. (2025) identify deficient theory of mind as a distinct inter-agent failure category in which agents fail to model each otherâs informational needs, and Kostka and Chudziak (2025) show that integrating Theory of Mind mechanisms improves multi-agent reasoning. ⢠Compound failure in iterative workflows. When an agent revisits an evolving artifact across sessions (e.g., multi-session debugging), failures compound: compacted memory does not distinguish original content from recent changes (missing Îť), discarded diff metadata prevents detecting what changed (destructive ĎâĎ^-), and the agent relies on stale memory rather than re-reading (Ď over Ď). This pattern arises naturally in any long-running workflow involving iterative revision. These transition-level failures manifest as operator omission or miscomposition rather than intrinsic zone pathologiesâeach is preventable using the existing operator set, and as the compound pattern shows, a single missing operator often triggers cascading failures in others. TableË1 summarizes the tripartite architecture. Table 1: The tripartite zonal architecture of contextual space. Zone-level failure modes are listed here; transition-level failures (contamination, destructive compaction, cross-agent incoherence, compound iterative failure) are characterized in the text. Zone AI Counterpart Cartographic Function Failure Mode Black Fog âŹB Tool use, search, code execution (Yao et al., 2023; Wang et al., 2023a) Reconnaissance, surveying Hallucination (Huang et al., 2023) Gray Fog G Structured memory, vector DBs (Jing et al., 2024), persisted state Archival mapping, state tracking Drift, staleness, bloat (Xiong et al., 2025) Visible Field V Active context window, prompt assembly (Khattab et al., 2024) Tactical projection surface Overload, attention dilution (Liu et al., 2024) 3.2 Salience Geometry of the Visible Field The visible field V is not a uniform surface. We characterize its geometry through a salience function that captures how information at different positions contributes to downstream reasoning. Definition 3 (Salience Function). For a context window of length n, the salience function s:1,âŚ,nĂââ[0,1]s:\1,âŚ,n\ĂNâ[0,1] assigns to each position i a weight reflecting its effective contribution to the modelâs next-token prediction. Under ideal (uniform) attention, sâ(i,n)=1/ns(i,n)=1/n for all i. Under empirical transformer attention (Vaswani et al., 2017), s exhibits a characteristic U-shaped profile. Salience Non-Uniformity. Liu et al. (2024) documented that s peaks at the beginning (primacy bias) and end (recency bias) of the sequence, with a trough in intermediate positions. Wu et al. (2025) provide a graph-theoretic proof that this non-uniformity arises from the interaction of causal masking and rotary positional encodings, and Li and others (2025) demonstrate a complementary âlost in the distanceâ effect: s degrades as the token distance between related elements increases, even when both remain within V. The problem compounds at scale. Vasylenko et al. (2025) show that normalized attention entropy grows with n as attention mass disperses across an expanding context, establishing a fundamental limitation: softmax-based transformers face increasing difficulty maintaining signal fidelity as context length grows. Capacity Constraints. The discrepancy between stated and effective context utilization is severe. Paulsen (2025) find that most models show significant accuracy degradation by 1,000 tokens, utilizing less than 1% of stated capacity; Wang and others (2026) identify critical thresholds beyond which performance collapses catastrophicallyâa phenomenon they term âshallow long-context adaptation.â These capacity limits are compounded by the rigid append-only paradigm of the KV cache (Li et al., 2025b): if an agent discovers during sense that a fact at position k is incorrect, it cannot update that belief without reprocessing the entire sequence from k onward. The discovery of âattention sinksâ (Xiao et al., 2024)âinitial tokens absorbing disproportionate attention regardless of contentâconfirms that positional gradients create structural anchors that further constrain effective capacity. Thesis 1 (Necessity of External Governance). Under the salience geometry characterized above, the visible field V cannot be treated as a passive container. Deliberate transformations are required to maintain reasoning coherence as |||V| grows. We term these transformations cartographic operators. Argument. The claim follows from the conjunction of four empirically established properties. (P1) Non-uniform s means that information placed at intermediate positions contributes less to downstream predictions than information at the extremes; a passive container offers no mechanism to ensure high-priority content occupies high-salience positions. (P2) Entropy growth means that as n increases, each additional token dilutes attention over all prior tokens; without active compression, signal-to-noise ratio degrades monotonically. (P3) The append-only KV cache means that incorrect or stale information at position k cannot be revised without reprocessing from k onward; a passive container has no mechanism to retract or update prior state. (P4) Effective capacity far below nominal means that the modelâs ability to utilize context degrades well before the stated window is filled; without active filtering, the agent operates on a fraction of the information it nominally has access to. Any system relying on passive concatenation inherits all four failure modes simultaneously; therefore, deliberate transformationsârepositioning, compression, eviction, filteringâare structurally required. Cartographic Laws of Context Representation. The necessity argument of Thesis 1 can be distilled into three principles that parallel foundational laws of cartographic representation (Snyder, 1987; MacEachren, 1995): 1. Bounded Surface Law. Reasoning requires a bounded visible field; not all information can be simultaneously active. This parallels the cartographic principle that maps are finite representations of unbounded territoryâthe cartographer must decide what to include, not merely how to render it. 2. Scale-Distortion Trade-off. Any compression of information into a bounded representation introduces distortion; the governance task is to choose which distortions are acceptable. In map projection, no single projection preserves angles, areas, and distances simultaneously (Snyder, 1987); in context management, no single transformation preserves all semantic properties simultaneously. 3. Salience Geometry Constraint. The bounded surface exhibits non-uniform effectiveness across positions; information placement affects reasoning quality. This parallels the cartographic principle that map readability depends on feature placement (visual hierarchy, figure-ground separation), not just feature presence. The seven cartographic operators defined below are the transformations that enforce these laws within and across epistemic zones. 3.3 Cartographic Operators Traditional cartography recognizes that a map is never a 1:1 reproduction of territoryâit is the result of deliberate operators applied to raw terrain (McMaster and Shea, 1992; Roth et al., 2011). We define seven operators for contextual space, derived from a systematic analysis of all non-trivial transformations within and between epistemic zones. For each, we specify: (a) a formal definition as a transformation, (b) the zones and transitions it governs, and (c) whether the correspondence with its classical cartographic counterpart is a structural isomorphism (preserving the same mathematical structure) or a functional analogy (serving the same purpose through different mechanisms). Design Principles. The operator set is organized by two orthogonal criteria: transformation type (what the operator does to information) and zone scope (where it applies). Most operators are zone-generalâthey define a transformation type that applies wherever information of the appropriate kind exists. One operator (δ, displacement) is zone-specific, arising from positional salience gradients present in sequential architectures. Three operators (Ď, Ď, and Ď) are boundary operators, governing how information crosses zone transitions: Ď determines where to explore, Ď determines which items cross a boundary, and Ď determines how items are represented upon crossing. For each boundary operator, the formal type signature (e.g., Ď:âŹââ(âŹ)Ď:B (B)) captures the decision the operator makes, while the zone scope in TableË3 indicates the transition it governs. Definition 4 (Reconnaissance). Ď:âŹââ(âŹ)Ď:B (B) determines which elements of unobserved territory to explore (what tools to invoke, what queries to issue, what code to execute). Governs sense transitions. The defining characteristic is decision under uncertainty: the agent does not know what the exploration will yield. Definition 5 (Selection). Ď:Zââ(Z)Ď:Z (Z) for Zâ,Zâ\V,G\ determines which elements cross a zone boundary. Governs recall, evict, and expire transitions via three modes: Ďrecall:ââ() _ recall:G (G) (select what to project into V), Ďevict:ââ() _ evict:V (V) (select what to archive from V), Ďexpire:ââ() _ expire:G (G) (enforce lifecycle policies). Formally, Ď is a policy over the source zoneâs inventory; its output parameterizes whichever transition it governs. Reconnaissance and selection are both boundary operators, but they differ in the epistemic status of their input. Ď operates under uncertainty: the agent does not know what âŹB contains and must plan exploration based on predictions about what might be useful. Ď operates over known inventory: the agent has metadata about G (or direct access to V) and can score items by relevance. This distinctionâplanner over unknown territory vs. filter over known contentâcorresponds to the explorationâexploitation boundary in decision theory (Cuconasu et al., 2024; Shi et al., 2023), and is tested directly in SectionË5.1. Definition 6 (Simplification). Ď:ZâZâ˛Ď:Zâ Z where |Zâ˛|<|Z||Z |<|Z| for zone Zâ,Zâ\V,G\, subject to the constraint that task-relevant semantic content of Z is preserved in Zâ˛Z . Zone-general. Definition 7 (Aggregation). Îą:z1,âŚ,zkâzâ˛Îą:\z_1,âŚ,z_k\â z fuses k semantically similar elements into a single representative, for elements in zone Zâ,Zâ\V,G\. Zone-general. Simplification reduces individual elements (e.g., compressing verbose tool output, condensing stored memory entries (Lindenbauer et al., 2025)). What constitutes âtask-relevant semantic contentâ depends on the downstream task, which is why simplification quality remains an empirical criterion rather than a formal guarantee. Aggregation merges multiple elements (e.g., deduplicating redundant observations, fusing repeated signals). Agentic memory systems implement aggregation explicitly (Xu et al., 2025), multi-agent episodic reconstruction (Wang et al., 2026) extends this across agent boundaries, and trajectory reduction methods (Xiao and others, 2025) illustrate the complementary role of selection (removing redundant steps rather than fusing them). Both simplification and aggregation apply within V and within G. Definition 8 (Projection). Ď:(Zsrc,p)âZdstĎ:(Z_src,p)â Z_dst transforms information across the âG boundary according to a projection schema p=(f,m,r,d)p=(f,m,r,d) specifying format f, modality m, resolution rârcoarse,âŚ,rfinerâ\r_coarse,âŚ,r_fine\, and structural dimensionality d. Bidirectional: Ď+:(,p)âĎ^+:(G,p) (forward projection) and Ďâ:(,p)âĎ^-:(V,p) (inverse projection, widely known as compaction). Forward projection subsumes operations that reduce structural dimensionalityâsuch as the context compression strategies of Kang et al. (2025) and the âstructure-then-selectâ approach of Zhou et al. (2025), which decomposes text into a discourse-unit tree and then selects query-relevant sub-trees for linearizationâand operations that adjust representational resolution (Liu and others, 2025; Chen et al., 2025). Both are parameters of the projection schema rather than independent operations. Every projection necessarily distorts: the choice of p determines which properties are preserved and which are sacrificed. A well-chosen projection maintains relational structure (adjacency, containment, causal links) even as the surface representation changes. A particularly severe distortion arises from modality mismatch: content stored in one modality but projected in another (e.g., a diagram rendered as text) may lose spatial structure even when propositional content is preserved. Definition 9 (Displacement). δ:(v,i)âŚ(v,j)δ:(v,i) (v,j) repositions element v from position i to position j in the sequence where sâ(j,n)>sâ(i,n)s(j,n)>s(i,n). Zone-specific (V only): compensates for the non-uniform salience function of SectionË3.2. Cartographic displacement repositions map features to resolve visual conflicts caused by finite spatial resolution; context displacement repositions information to compensate for non-uniform attention salience (Liu et al., 2024; Wu et al., 2025; Hsieh et al., 2024). Both sacrifice positional accuracy for functional effectiveness, but compensate for different medium constraints (2D overlap vs. 1D salience gradients). Common implementation patterns include constraint pinning (re-projecting invariant constraints into the first k tokens at each turn), recency injection (appending a rolling summary of high-priority state to the sequence end), and salience-aware assembly (ordering segments to align critical content with primacy and recency peaks). Displacement is irreducible to selection: Ď decides which items enter V; δ decides where they sit once inside. Definition 10 (Layering). Îť:ZâZ1ĂâŻĂZmÎť:Zâ Z_1Ă¡sĂ Z_m partitions a zone Zâ,Zâ\V,G\ into m typed namespaces. Zone-general. Within V, layering separates system constraints, task state, retrieved memory, and fresh observations into distinct semantic layers with explicit priority orderingâWallace et al. (2024) demonstrate that training LLMs to respect a typed instruction hierarchy (system >> user >> third-party) dramatically improves robustness to prompt injection. Within G, layering organizes stored memory into typed containers (as in MemOSâs MemCube abstraction (Li et al., 2025d)). Layering can be applied recursivelyâeach namespace itself partitioned into sub-namespacesâproducing the spatial hierarchies fundamental to both cartography and context engineering (e.g., repo â package â module â class â method in a codebase, or domain â concept â fact in a knowledge base). Recursive Îť is the formal mechanism underlying hierarchical G structures such as OpenVikingâs filesystem tree (Volcengine, 2025) and MemOSâs nested graph schemas (Li et al., 2025e). 3.4 Operator Properties and Cartographic Correspondences Each operator corresponds to a classical cartographic counterpart. We now characterize these correspondences, the invariants they preserve, and the points at which the geographic and contextual settings diverge. Cartographic Correspondences. Each operator has a counterpart in cartographic generalization, classified as either a structural isomorphism or a functional analogy. The classification criterion is: a correspondence is an isomorphism when the cartographic and context operators share input/output type structure and composition behavior (so that theorems about one transfer to the other), and an analogy when they serve the same functional role through structurally different mechanisms (so that design intuitions transfer but formal results do not). Ď, Îą, Ď, and Îť are structural isomorphisms: Ď preserves relevance ranking, Îą preserves semantic equivalence classes, Ď preserves relational structure across representation change, and Îť preserves namespace disjointness. ĎĎ and δ are functional analogies: cartographic simplification reduces geometric complexity while context simplification reduces token count (Lindenbauer et al., 2025); cartographic displacement resolves visual overlap while context displacement compensates for salience gradients. TableË3 records the correspondence type for each operator; TableË2 provides the detailed mapping from classical cartographic operators to their context counterparts. Representational Invariants. In classical cartography, map projections are classified by the property they preserve: conformality (local shape), equivalence (area), or equidistance (distance) (Snyder, 1987). A foundational result is that these are mutually exclusiveâno projection preserves all three simultaneously. We identify an analogous triad for context transformations: ⢠Semantic invariance: the meaning of individual elements is preserved despite surface-level reduction. Primarily enforced by ĎĎ (simplification). ⢠Structural invariance: relational structure (adjacency, containment, causal links) is preserved despite representational change. Primarily enforced by Ď (projection). ⢠Salience invariance: reasoning relevance is preserved despite positional change. Primarily enforced by δ (displacement). As with map projections, no single context transformation preserves all three invariants simultaneously: simplification sacrifices structural detail, projection sacrifices surface semantics, and displacement sacrifices chronological ordering. The existence of multiple operators reflects this fundamental trade-off. Distinguishing ĎĎ, Îą, and Ď. An apparent redundancy deserves explicit justification: ĎĎ (simplification, DefinitionË6), Îą (aggregation, DefinitionË7), and Ď (projection, DefinitionË8) all transform representations, and one might argue they should be unified into a single âtransformationâ operator. We maintain the three-way split because each preserves a distinct invariant, fails in a distinct mode, and corresponds to a distinct classical cartographic operator: ⢠ĎĎ preserves semantic content while reducing surface tokens. Input: one element. Output: one shorter element. Failure mode: critical details dropped (Zhang et al., 2025b; Lindenbauer et al., 2025). Cartographic analogue: simplifying a coastlineâs vertex count while preserving its shape (Roth et al., 2011). ⢠ι preserves the equivalence class while reducing element count. Input: k similar elements. Output: one composite representative. Failure mode: distinguishing features between merged elements lost. Cartographic analogue: merging individual buildings into a city-block polygon (McMaster and Shea, 1992). â˘ Ď preserves relational structure while changing the representational system. Input: one element in format A. Output: one element in format B. Failure mode: structural information lost in format translation (Zhou et al., 2025). Cartographic analogue: projecting coordinates from sphere to plane. Collapsing these into a single âtransformationâ operator would prevent diagnosing which specific invariant was violated when a transformation fails. Table 2: Mapping from classical cartographic generalization operators to context cartography operators. âIsomorphismâ indicates shared formal structure; âAnalogyâ indicates shared purpose with different mechanism; âSubsumedâ indicates the cartographic operator maps to a mode or parameter of a context operator rather than a standalone counterpart. Cartographic Op. Classical Function Context Op. Corr. Notes Selection / elimination Filter features by relevance Ď Iso. Both filter known inventory by task criteria Simplification Reduce geometric complexity ĎĎ Ana. Cartographic: vertex count; Context: token count Aggregation Merge similar features Îą Iso. Both fuse proximate items into composites Typification Replace dense set with representatives Ď (mode) Sub. Pattern-preserving sampling; a selection strategy Displacement Resolve visual conflicts δ Ana. Cartographic: 2D overlap; Context: 1D salience Collapse Reduce dimensionality Ď (param.) Sub. Dimensionality reduction within projection schema Symbolization Assign typed visual categories Îť Iso. Both enforce semantic namespace separation Map projection Transform between coordinate systems Ď Iso. Both transform representations with controlled distortion Context operators without classical cartographic generalization counterpart: Reconnaissance (Field surveying / data acquisition) Ď â Precedes generalization; cartographic analogue is survey planning Where Correspondences Diverge. The isomorphisms and analogies in TableË2 hold at the level of formal structure, but the underlying media differ in ways that affect operator behavior. Three systematic divergences are worth noting. First, cartography operates on continuous 2D space with visual perception constraints, while context operates on discrete 1D sequences with attention constraints; this means cartographic displacement resolves pairwise feature conflicts (a local, geometric problem), whereas context displacement compensates for a global salience gradient imposed by the architecture. Second, cartographic generalization is typically performed offline during map compilation, while context operators must execute online within the latency budget of each agent turnâimposing computational constraints that classical cartographic operators do not face. Third, geographic features have stable identities across scales (a river is the same river at 1:50,000 and 1:1,000,000), whereas context elements may lose identity through compaction: a summarized conversation turn is not the âsameâ element at a different resolution but a new synthetic artifact. These divergences do not invalidate the correspondences but constrain how directly cartographic algorithms can be transferred to context engineering. 3.5 Composition and Architecture We now show that the seven operators cover all zonal transformations, compose into pipelines, and integrate into a minimal reference architecture. Coverage. The seven operators cover all non-trivial transformations across the zonal architecture: ⢠âŹâB : governed by Ď (reconnaissance) ⢠âG : governed by Ď (selection) ++ Ď+Ď^+ (forward projection) ⢠âV : governed by ĎĎ, Îą, δ, Îť ⢠âV : governed by Ď (select what to evict) ++ ĎâĎ^- (compaction) ⢠âG : governed by ĎĎ, Îą, Îť applied within G (memory consolidation) ⢠ââŹG : governed by Ď (lifecycle-aware expiration policy) This coverage claim, along with the zonal partition and layering disjointness properties, has been mechanically verified in Lean 4 (AppendixËB). Completeness. Several prima facie gaps in the operator set reduce, on analysis, to failure modes of existing operators or their compositions rather than to missing primitives: ⢠Output gating (mediating raw sense output before it enters V) is not a separate operator but the prescribed application of ĎâĎ+Ď Ď^+ on the inbound path; its absence produces visible field contamination. ⢠Multi-agent context packaging (an orchestrator composing a context payload for a subagent) is a composition ÎťâĎ+âĎÎť Ď^+ Ď applied to the orchestratorâs own V or G, with the subagentâs initial context as destination. Poor packaging is a failure of Ď (wrong selection), Ď+Ď^+ (wrong format), or Îť (missing namespace separation). ⢠Faithfulness verification (checking that a transformation preserved task-relevant semantics) is a quality criterion for operators, not a transformation itselfâanalogous to map validation in cartography, which evaluates generalization operators but is not itself a generalization operator. ⢠In-place context reset (replacing an entire V with a condensed summary, as in conversation compaction) decomposes into a compaction cycle: ĎâĎ^- (archive original to G) followed by ĎâĎ+Ď Ď^+ (select and project summary back into V). We reserve compaction for the ĎâĎ^- step alone (not the full cycle). When the archival step is skipped, the original is lost to âŹBâproducing context collapse. Composition. Operators compose into cartographic pipelines. The inbound pipeline from gray fog to visible field applies Ď (select what to recall), then Ď+Ď^+ (project into token format), then ĎĎ (simplify), then δ (position for salience), then Îť (assign to namespace): ÎťâδâĎâĎ+âĎÎť δ Ď Ď^+ Ď. The outbound pipeline reverses this: Ď (select what to evict), then ĎâĎ^- (compact into storage format): ĎââĎĎ^- Ď. A maintenance cycle applies Ď,Îą,Îť\Ď,Îą,Îť\ within G asynchronouslyâthe formal characterization of processes such as Lettaâs sleep-time memory reorganization (Letta, 2025). Two basic composition properties are worth noting. First, ĎĎ is approximately idempotent: applying simplification twice yields marginal further reduction, since most redundancy is removed in the first pass. Second, δ and ĎĎ do not commute in general: simplifying before repositioning (δâĎδ Ď) operates on already-compressed content, while repositioning before simplifying (ĎâÎ´Ď Î´) may discard content that was just moved to a high-salience position. The prescribed pipeline order places ĎĎ before δ to avoid this interaction. A systematic analysis of commutativity and idempotency across all operator pairs remains an open problem (SectionË5.3). FigureË2 illustrates the inbound, outbound, and maintenance pipelines. GVĎĎ+Ď^+ĎĎδΝ (recall)ĎĎâĎ^-Outbound (evict)Ď,Îą,ÎťĎ,Îą,Îť Figure 2: Cartographic pipelines. The inbound pipeline ÎťâδâĎâĎ+âĎÎť δ Ď Ď^+ Ď transforms content from G to V; the outbound pipeline ĎââĎĎ^- Ď archives from V to G; the dashed loop represents asynchronous maintenance within G. Scale Modulation. In cartography, map scale is the ratio between a distance on the map and the corresponding distance on the ground; it determines which features are representable and which operators apply. We define contextual scale analogously as the ratio between the information content of a source entity and its representation in V. At coarse contextual scale, a 500-line module is represented by a one-sentence summary; at fine contextual scale, the full source is projected. Contextual scale is not a single operator but a coordination policy that adjusts all operator parameters simultaneously (McMaster and Shea, 1992; Roth et al., 2011): reducing the scale increases simplification intensity, triggers aggregation, broadens selection scope, and may suppress entire layers. Prior work on multi-resolution context managementâincluding hierarchical compression (Kang et al., 2025), discourse-unit decomposition (Zhou et al., 2025), and tiered memory loading (Volcengine, 2025)âaddresses individual resolution choices within a single operator. Each âzoom level,â however, does not merely change Ďâs resolution parameter râit changes the parameters of Ď, ĎĎ, Îą, and Îť in concert. Scale modulation is therefore not an eighth operator but a policy over the pipelineâs parameterization, governing how operator settings co-vary with a target resolution. In practice, the resolution levels available to this policy correspond to levels in a spatial hierarchy: zooming in descends the hierarchy, zooming out ascends it. Aggregation (Îą) constructs these hierarchiesâmerging sibling elements into their parent levelâwhile scale modulation navigates them. Among the systems in SectionË4, OpenViking and Letta implement explicit scale policies (tiered loading and core/recall separation, respectively), while Claude Code makes per-operation resolution choices without a coordinating policy. Minimal Reference Architecture. A cartographically governed agent maintains a structured G and a context assembler that builds V at each turn by applying the inbound pipeline, the outbound pipeline, and asynchronous G-maintenance as defined above. A decision layer at the Ď/Ď boundary routes knowledge gaps to tool invocation or memory recall based on the agentâs epistemic state. TableË3 summarizes the seven operators. Table 3: Cartographic operators for contextual space transformation. Operator Formal Role Zone Scope Cart. Corr. Reconnaissance Ď Determine what to explore âŹâB â Selection Ď Filter known inventory by relevance âG Isomorphism Simplification ĎĎ Reduce tokens, preserve structure V or G Analogy Aggregation Îą Fuse repeated signals V or G Isomorphism Projection Ď Transform representation across boundary âG Isomorphism Displacement δ Reposition for salience V only Analogy Layering Îť Separate into typed namespaces V or G Isomorphism Mapping to Engineering Vocabulary. TableË4 connects the framework to common agent engineering concepts, showing which zones, transitions, and operators each concept involves. Table 4: Common agent engineering concepts mapped to the cartographic framework. Concept Zone / Transition Operators Tool / function call sense: âŹâB Ď (decide to call); ĎâĎ+Ď Ď^+ (mediate result before entering V) RAG retrieval recall: âG Ď (query indexed corpus); Ď+Ď^+ (format chunks for V) MCP server (Hou et al., 2025) Standardized sense interface Ď (select server/tool); Ď (common projection schema across providers) System prompt (Mu et al., 2025) Pinned region of V Îť (typed namespace); δ (constraint pinning at primacy position) Conversation history (Wang et al., 2023b) V (live); G (after compaction) ĎâĎ^- (compaction cycle); Ďrecall _ recall (re-project from archive) Memory system (Sumers et al., 2023; Pink et al., 2025) G Îť (episodic / semantic / procedural layers); ĎĎ, Îą (maintenance) Subagent delegation (Zhang et al., 2025c) âŹB from orchestratorâs view Ď (delegate exploration); Îť (package context); ĎâĎ+Ď Ď^+ (receive condensed result) Summarization âV or âV ĎĎ (within-zone); ĎâĎ^- (cross-zone compaction) Context window management Cycle: âV Full compaction cycle: ĎâĎâĎ Ď^- (archive), then ĎâĎ+Ď Ď^+ (re-project summary) 4 Case Studies: Emergent Cartographic Practices We analyze four contemporary systems to demonstrate the descriptive utility of the frameworkâits ability to capture, in a unified vocabulary, design patterns that these systems developed independently for different use cases. To ensure reproducibility, we evaluate each systemâoperator pair against a five-criterion rubric measuring implementation depth (SectionË4.5). An important caveat: the operators in SectionË3.3 were informed in part by observing these systems, so the case studies are not an independent test of the frameworkâs predictive power. They demonstrate that the vocabulary is descriptively adequateâthat it can account for the design choices these systems makeâbut not that it predicted those choices. Predictive validation requires testing against systems not used during framework development, which we identify as a priority for future work. 4.1 Claude Code: Subagent Isolation Claude Code (Anthropic, 2025) implements cartographic zoning: it spawns subagents with bounded exploration contexts, separating sense operations from the primary V. The orchestrator delegates Ď (reconnaissance) to subagents, which traverse âŹB in isolation, then apply ĎĎ (simplification) and ĎâĎ^- (compaction) before the orchestrator projects condensed results back into its visible field via Ď+Ď^+. This achieves a weaker form of δ (displacement) through salience-advantaged placement rather than active repositioning: summaries enter at the recency-biased end of the sequence, benefiting from positional salience without explicit governance. The primary gap is explicit resolution control within Ďâsubagent summaries are returned at a single projection schema. The key architectural contribution is the agentic handoff: the orchestrator must compose a context package for each subagent via Îť (layering), transferring task-relevant constraints without projecting the entire global state. 4.2 Letta (MemGPT): Memory as Operating System Letta (Packer et al., 2023) restructures G as an operating-system memory hierarchy. Core memory blocks are perpetually projected into V via fixed Ď+Ď^+ (implementing persistent recall), while recall memory manages overflow through recursive ĎâĎ^- (compaction)âevicted content is condensed and archived. Background âsleep-timeâ processes implement G-internal maintenance by applying Ď,Îą,Îť\Ď,Îą,Îť\ within G, periodically reorganizing stored memory into hierarchical structures. Context repositories (Letta, 2025) extend G with git-based versioning for persistent state. The core/recall tier separation implements resolution control within Ď: core memory is projected at full resolution, recall memory at reduced resolution. 4.3 MemOS: Graph-Structured Memory MemOS (Li et al., 2025d, e) treats G as enterprise-grade infrastructure. The MemCube abstraction implements Îť (layering) within G by isolating distinct knowledge bases as composable, typed containers. DAG-based scheduling implements Ď (selection) across multi-stage workflows. The ânext-scene predictionâ mechanism (Li et al., 2025e) implements predictive Ď and Ď+Ď^+âpreloading memory at appropriate resolution before explicit request. The graph-structured storage enables rich Ď+Ď^+ projection schemas: the same underlying memory can be projected as entity lists, relationship summaries, or full subgraph traversals depending on the task. Empirical results show a 159% gain in temporal reasoning and 60.95% token reduction on the LoCoMo benchmark (self-reported). 4.4 OpenViking: Hierarchical Context OpenViking (Volcengine, 2025) spatializes G as a hierarchical filesystem with a custom URI scheme. Its tiered loading protocol (L0: âź 100 tokens, L1: âź 1,000 tokens, L2: full content) is a direct implementation of multi-resolution Ď+Ď^+ (forward projection with explicit resolution parameter). Directory recursion implements structured Ď (selection) via path-based traversal rather than flat similarity search. The system preserves and visualizes the âretrieval trajectory,â implementing a structural reduction within Ď+Ď^+ by presenting the navigated path rather than the full filesystem. 4.5 Cross-System Analysis To move beyond binary present/absent judgments, we evaluate each systemâoperator pair on five criteria, each scored 0 or 1: 1. Present (P): any mechanism performs this transformation. 2. Explicit (E): the mechanism is a named architectural component, not an emergent side-effect. 3. Configurable (C): the operatorâs parameters can be tuned by the user or developer. 4. Automated (A): the transformation occurs without manual intervention during normal operation. 5. Documented (D): the mechanism is described as a feature in official documentation or papers. Each cell in TableË5 reports a score from 0 (absent) to 5 (fully realized). Assignments are evidence-based lower bounds: they reflect what can be determined from public documentation, technical papers, and source code (where available), and may understate internal capabilities not externally visible. A key scoring distinction: emergent side-effects (e.g., content entering at the recency-biased end of an append-only sequence) may satisfy P (Present) but not E (Explicit), since no deliberate architectural mechanism performs the transformation. This is why δ scores remain low despite every append-only system exhibiting incidental positional bias. Table 5: Operator implementation depth across four systems. Each cell scores five binary criteria (Present, Explicit, Configurable, Automated, Documented; max = 5). Row means indicate per-operator adoption; column means indicate per-system coverage. Operator Claude Code Letta MemOS OpenViking Mean Reconnaissance Ď 5 1 1 1 2.00 Selection Ď 2 2 5 5 3.50 Simplification ĎĎ 4 4 2 1 2.75 Aggregation Îą 1 2 4 1 2.00 Projection Ď 2 5 5 5 4.25 Displacement δ 2 1 1 1 1.25 Layering Îť 5 5 5 5 5.00 System mean 3.00 2.86 3.29 2.71 2.96 Four quantitative observations emerge from the scored analysis. First, Îť (layering) achieves a perfect mean score of 5.00âit is the only operator fully realized as a primary, configurable, automated, and documented mechanism in every system. This confirms that namespace contamination is the most universally recognized failure mode (Zhang et al., 2025b). Second, Ď (projection) scores 4.25, the highest among non-universal operators, indicating strong convergence on the need for representational transformation at the âG boundary. Third, δ (displacement) scores only 1.25âthe lowest of any operatorâdespite addressing the best-documented failure mode (Liu et al., 2024). Where it exists, it is implicit (e.g., summaries entering at the recency-biased end of the sequence in Claude Code) rather than an explicit architectural mechanism. Fourth, the systems bifurcate along the Ď/Ď axis: Claude Code emphasizes exploration governance (Ď=5Ď=5) while MemOS and OpenViking emphasize projection governance (Ď=5Ď=5, Ď=5Ď=5). Letta bridges both strategies with strong projection (Ď=5Ď=5) and moderate simplification (Ď=4Ď=4). The overall system mean of 2.96 out of 5.00 indicates that current systems implement approximately 60% of the cartographic operator space, with systematic gaps in reconnaissance (Ď), aggregation (Îą), and displacement (δ). The transition-level failure modes of SectionË3.1 map onto these gaps: low ĎĎ scores expose systems to visible field contamination, low Ď scores leave the explorationâexploitation boundary ungoverned, and the universal adoption of Îť (score = 5 across all systems) reinforces the namespace contamination finding above. 5 Research Agenda The preceding framework makes testable claims: that cartographic operators improve reasoning coherence, and that their effects are attributable to compensating for specific geometric properties of the visible field. This section derives falsifiable predictions from the framework and proposes a diagnostic benchmark for community evaluation. 5.1 Proposed Benchmark: Context Cartography Diagnostic (CCD) We propose a diagnostic benchmark with task categories designed to isolate individual operators and zone transitions. Unlike general-purpose long-context benchmarks (Paulsen, 2025), the CCD is designed to measure operator-specific effects through targeted task design: ⢠Reconnaissance vs. selection tasks. The agent must decide when to invoke tools (Ď: explore âŹB) versus answering from existing memory (Ď: recall from G), measuring the explorationâexploitation boundary. The agent is penalized both for unnecessary tool calls (wasteful Ď) and for hallucinating answers that could have been grounded by Ď. ⢠Projection tasks. The agent receives information in various storage-native formats (JSON, graph structures, hierarchical file listings) and must project it into effective reasoning context. Tasks require switching between summary-level and detail-level reasoning within a single session (e.g., first identify the relevant module at coarse resolution, then debug a specific function at fine resolution), testing multi-resolution Ď. ⢠Displacement tasks. Critical constraints (safety rules, formatting requirements) are placed at varying positions within the context. The agent must adhere to these constraints regardless of position. This extends the GM-Extract protocol (Gupte et al., 2025) with agentic evaluation. ⢠Simplification tasks. The agent receives either raw tool output (verbose logs, full code listings) or simplified summaries, and must complete a downstream task. This measures whether ĎĎ preserves task-relevant information. ⢠Aggregation tasks. The agent receives k partially overlapping observations (e.g., multiple search results covering the same topic, or repeated tool outputs with minor variations) and must produce a unified answer. Without Îą, redundancy accumulates in V, consuming token budget and diluting attention across repeated content. ⢠Layering tasks. Conflicting information is placed in different semantic layers (system prompt vs. retrieved memory vs. user input). The agent must correctly prioritize based on layer type. The benchmark is designed for evaluation via operator ablation: implementing the cartographic pipeline as composable middleware layers and systematically removing individual operators to measure their isolated contribution. Each configuration would be evaluated on task accuracy, token consumption, and failure mode distribution (hallucination rate, constraint adherence, information loss). Failure mode labeling can be instrumented through tool-call logging (detecting unnecessary Ď), gold-provenance comparison (distinguishing hallucination from stale-memory reliance), and constraint-checker oracles (measuring δ/Îť adherence). 5.2 Testable Predictions The framework, combined with the transition-level failure analysis (SectionË3.1), yields specific falsifiable predictions. Each prediction identifies an operator, the failure mode expected when that operator is absent, and the observable effect: 1. Simplification necessity (ĎĎ). Removing simplification from the inbound pipeline increases visible field contamination: raw tool output displaces task-relevant content, degrading downstream accuracy. The effect should scale with tool output verbosityâhigh-verbosity tasks (full code listings, build logs) should show larger degradation than low-verbosity tasks (structured API responses). 2. Displacementâlength interaction (δ). Removing salience-aware positioning increases constraint violation rates, with the effect size scaling with context length as the salience trough deepens (Liu et al., 2024). At short context lengths (below âź 4K tokens), displacement should have negligible effect; at long context lengths (>>32K tokens), the effect should be substantial. 3. Layeringâconflict interaction (Îť). Removing namespace separation increases layer-priority errors specifically when information sources conflict. On tasks with no inter-source conflicts, layering should have minimal effect; on tasks with deliberate conflicts (e.g., system prompt contradicts retrieved memory), the effect should be large. 4. Compaction mode (ĎâĎ^-). Under archival compaction (ĎâĎ^- with original preserved in G), information loss should remain bounded over successive compaction cycles. Under destructive compaction (ĎâĎ^- without archival), information loss should accumulate, with the gap widening over iterationsâthe quantitative signature of context collapse (Zhang et al., 2025b). 5. Explorationâexploitation boundary (Ď/Ď). Agents without explicit reconnaissance governance should exhibit a bimodal failure distribution: over-reliance on G (hallucination from stale memory when Ď is under-used) or over-exploration of âŹB (wasteful tool calls when Ď is under-used). Explicit Ď/Ď governance should reduce variance in exploration behavior. 5.3 Open Problems The framework raises several questions that go beyond what the current formalism can answer: ⢠Optimal composition order. The inbound pipeline ÎťâδâĎâĎ+âĎÎť δ Ď Ď^+ Ď prescribes a fixed order. Are there task-dependent orderings that improve performance? For instance, should δ precede ĎĎ when the simplification itself is position-sensitive? ⢠Operator interaction effects. The predictions above treat operators as independent. In practice, operators may interact: strong Îť (layering) might partially compensate for weak δ (displacement) by structurally isolating high-priority content. Understanding these interactions requires factorial experimental designs beyond single-operator ablation. ⢠Cross-architecture transfer. The salience geometry of SectionË3.2 is specific to softmax-based transformers. State space models and diffusion language models (SectionË6.1) may exhibit different salience profiles, potentially rendering δ unnecessary while creating demand for new operators. The completeness argument (SectionË3.3) holds for the current transformer-based zonal structure; non-transformer architectures may alter that structure and with it the operator basis. ⢠Automated operator selection. Can a meta-controller learn which operators to apply and with what parameters, given task characteristics and current context state? This connects the framework to the broader question of learned context management (Packer et al., 2023; Li et al., 2025e). 6 Discussion 6.1 Which Operators Survive Architecture Change? The seven operators compensate for specific properties of transformer architectures: linear prefix memory, append-only KV caches, and entropy growth under softmax. Emerging architectures relax different subsets of these constraints, and the framework predicts which operators each architecture renders obsolete versus which it preserves. State Space Models (Gu and Dao, 2024; Dao and Gu, 2024) and their production-scale hybrids (Blakeman et al., 2025; Lieber et al., 2024; Glorioso et al., 2024)âas well as pure recurrent architectures such as RWKV-7 (Peng and others, 2025)âcompress history into evolving latent states with constant-memory processing. This eliminates the append-only constraint, potentially internalizing δ (displacement): if the model can revise its state representation in-place, positional salience gradients no longer dictate where information must be placed. However, the compression is lossyâthe latent state is itself a form of ĎâĎ^- (compaction) applied continuously and without external governance. Conjecture 1 (Implicit Context Collapse). SSM-based agents will exhibit implicit context collapse: information loss that accumulates inside the latent state without any mechanism to detect or reverse itâanalogous to repeated ĎâĎ^- without archival (SectionË3.1), but executed continuously within the modelâs hidden state rather than as a discrete external step. If Ë1 holds, then ĎĎ (simplification) and Ďexpire _ expire (lifecycle-aware selection) become more important under SSMs, not less, because the compaction is no longer an explicit step that can be audited. Diffusion language models (Nie et al., 2025; Arriola et al., 2025; Ye et al., 2025) weaken the left-to-right constraint, enabling local non-causal refinement. This relaxes the salience geometry of SectionË3.2: if the model can attend bidirectionally, the U-shaped salience profile flattens, and δ loses its rationale entirely. But diffusion models introduce a new constraintâiterative denoising across the full sequenceâthat creates computational pressure to keep |||V| small. The framework predicts that ĎĎ and Îą (aggregation) will be critical for diffusion-based agents, not to compensate for attention non-uniformity but to keep the denoising target tractable. Architectures that treat context as an external navigable environment (Zhang et al., 2025a) or that modify weights at test time (Behrouz et al., 2025; Zweiger et al., 2025) push further: they begin to internalize G itself, blurring the boundary between stored memory and model parameters. Under such architectures, Ď (projection) shifts from an external engineering task to an internal learned capability. Yet the zonal structure persists: the world will always exceed any modelâs capacity (âŹâ â Bâ ), memory will always require governance (G cannot be unbounded), and reasoning will always require a bounded surface (|||V| is finite). What changes is not the need for cartographic operators but where they executeâexternally as middleware, or internally as learned model behavior. 6.2 Cartographic Competencies as Training Objectives Current LLMs are optimized to continue text, not to manage contextual space. Yet agent behavior demands cartographic competencies: choosing between Ď (reconnaissance) and Ď (selection) when facing knowledge gaps, applying Ď and ĎĎ to recall outputs, and respecting Îť constraints in V. The framework identifies three specific mismatches between next-token prediction training and cartographic requirements. First, next-token prediction treats all positions as equally important prediction targets, but the salience geometry of SectionË3.2 shows that not all positions contribute equally to downstream reasoning. Training paradigms that weight prediction targets by information densityâsuch as patch-level training (Shao et al., 2025) and mutual-information-aware objectives (Yang and others, 2025)âimplicitly learn a form of ĎĎ (simplification): the model learns which tokens carry structural signal and which are surface syntax. Second, the Ď/Ď boundary (when to explore âŹB vs. exploit G) is not present in pre-training at all. Standard training never requires the model to decide âI donât know thisâI should invoke a toolâ versus âI saw this earlierâI should recall it.â This competency must currently be elicited through prompting or reinforcement learning on agentic trajectories. The framework suggests that training explicitly on Ď/Ď decision pointsâtasks where the optimal action depends on the agentâs epistemic state across zonesâwould improve tool-use calibration. Third, surveys of System 1 to System 2 reasoning in LLMs (Li et al., 2025f; Kahneman, 2011) catalog methods (MCTS, RL fine-tuning, macro actions) that require the agent to maintain structured state. Li et al. (2025a) provide direct evidence that the structure of reasoning demonstrationsânot their contentâdetermines reasoning quality: disrupting the logical organization of chain-of-thought steps collapses accuracy, while substituting incorrect content has minimal effect. This supports our argument that deliberate reasoning (System 2) is mediated by how context is structured: it requires the agent to maintain layered state (Îť), track what has been explored (Ď) versus what is assumed (Ď), and revise beliefs when new evidence arrives. These are cartographic competencies, and their absence under standard training explains why scaling model size alone does not reliably produce agentic capability. 6.3 Multi-Agent Cartography Multi-agent systems (Tran et al., 2025) extend the framework from a single zonal structure to a federation of maps. Each agent i maintains its own (âŹi,i,i)(B_i,G_i,V_i), and agents may share a common memory sharedG_ shared with layered provenance (Îť applied across agent boundaries). Cross-agent communication is modeled as constrained Ď+Ď^+: agent i projects a result from iV_i into sharedG_ shared (or directly into jV_j) using a projection schema that both parties must agree on. The central question is how these maps interact. The framework identifies three distinct multi-agent failure modes, each corresponding to a specific operator failure across agent boundaries: ⢠Projection incompatibility. Agent A applies Ď+Ď^+ with schema pAp_A to communicate a finding; agent B receives it but interprets it under schema pBp_B. If pAâ pBp_Aâ p_B, the receiving agentâs V contains well-formed but semantically distorted contentâthe cross-agent projection incoherence identified in SectionË3.1. ⢠Namespace collision. Two agents write to a shared G without coordinated Îť (layering). Their contributions are interleaved without provenance marking, making it impossible for a third agent to distinguish or prioritize between them. This is the multi-agent analogue of the single-agent layering failure. ⢠Reconnaissance duplication. Multiple agents independently apply Ď to the same region of âŹB, wasting exploration budget. Without a shared record of what has been explored, the system cannot distinguish âŹB (unexplored) from G (explored by another agent but not yet shared). Global Workspace Theory (Baars, 1988; Goldstein and Kirk-Giannini, 2024) offers an architectural response: a governed broadcast bottleneckâanalogous to a shared Vâthrough which all inter-agent communication must pass. In cartographic terms, this shared workspace forces agents to use a common projection schema for cross-agent Ď+Ď^+, a coordinated namespace for shared G via Îť, and a joint exploration ledger for Ď. Empirical evidence supports this: Yuen et al. (2025) demonstrate that agent-specific structured memories shared across heterogeneous agents achieve state-of-the-art coordination, and cognitive-workspace frameworks (An, 2025; Li et al., 2025c) show that structured broadcast improves coherence even without specialized trainingâconsistent with the prediction that the benefit comes from operator coordination rather than model capability. 7 Limitations Several limitations of this work should be acknowledged. First, the framework is primarily interpretive: the cartographic operators are derived from cartographic generalization theory, and while we distinguish structural isomorphisms from functional analogies (SectionË3.3), the formalization does not yet constitute a full mathematical theory. Second, the case study analysis (SectionË4) is post-hocâthe operators were informed in part by observing these systems, so the convergence finding is interpretive rather than predictive. Third, while the operators are derived from systematic coverage of zone transformations, additional cartographic operators (e.g., symbolization, rotation) may have context analogues not yet identified, and the zone-general scope of ĎĎ, Îą, and Îť may warrant finer distinctions as empirical evidence accumulates. Fourth, the research agenda (SectionË5) derives testable predictions but does not empirically validate them; the proposed CCD benchmark requires community implementation and evaluation. Fifth, the zonal model is defined for a single agent; the multi-agent extension (SectionË6.3) identifies failure modes and architectural responses but does not formally extend the zone definitions or operator signatures to federated settings. Finally, the extension of cartographic theory from geographic to contextual space introduces representational differences: contextual space is sequential and high-dimensional, while geographic space is continuous and low-dimensional. The extent to which cartographic principles transfer across this gapâand whether the invariant triad (semantic, structural, salience) fully characterizes the contextual distortion spaceârequires further investigation. 8 Conclusion We have presented Context Cartography, a formal framework for governing contextual space in LLM systems. The framework contributes a tripartite zonal model with explicit state transitions, a characterization of salience geometry grounding governance strategies in attention physics, and seven formally defined cartographic operatorsâreconnaissance, selection, simplification, aggregation, projection, displacement, and layeringâderived from systematic coverage of all zone transformations and organized by transformation type and zone scope. Analysis of four contemporary systems provides interpretive evidence that these operators are converging independently across the industry. The central claim of this work is that contextual space has geometry, and that intelligence depends on how that geometry is governed. Current cartographic practices are compensatoryâthey exist because no architecture natively supports the full range of spatial governance that agentic reasoning demands. As architectures evolve, some operators may internalize (transitioning from external engineering to learned model behavior), but the zonal structure persists: the world exceeds any modelâs capacity, memory requires governance, and reasoning requires a bounded surface. The testable predictions and diagnostic benchmark proposed in our research agenda provide a path toward empirical validation. References T. An (2025) Cognitive workspace: active memory management for LLMs â an empirical study of functional infinite context. arXiv preprint arXiv:2508.13171. Cited by: §6.3. Anthropic (2025) Claude code documentation. Note: https://code.claude.com/docs/en/overviewAccessed: 2026-02-27 Cited by: item Ď (Reconnaissance) = 5 (P,E,C,A,D), §4.1. M. Arriola, A. Gokaslan, J. T. Chiu, Z. Yang, Z. Qi, J. Han, S. S. Sahoo, and V. Kuleshov (2025) Block diffusion: interpolating between autoregressive and diffusion language models. In International Conference on Learning Representations, Note: Oral presentation Cited by: §6.1. B. J. Baars (1988) A cognitive theory of consciousness. Cambridge University Press. Cited by: §6.3. A. Baddeley (2000) The episodic buffer: a new component of working memory?. Trends in Cognitive Sciences 4 (11), p. 417â423. Cited by: §3.1. A. Behrouz, P. Zhong, and V. Mirrokni (2025) Titans: learning to memorize at test time. In Advances in Neural Information Processing Systems, Cited by: §6.1. I. Beltagy, M. E. Peters, and A. Cohan (2020) Longformer: the long-document transformer. arXiv preprint arXiv:2004.05150. Cited by: §2. A. Blakeman, A. Basant, A. Khattar, A. Renduchintala, A. Bercovich, A. Ficek, et al. (2025) Nemotron-h: a family of accurate and efficient hybrid mamba-transformer models. arXiv preprint arXiv:2504.03624. Cited by: §6.1. K. E. Brassel and R. Weibel (1988) A review and conceptual framework of automated map generalization. International Journal of Geographical Information Systems 2 (3), p. 229â244. Cited by: §2. M. Cemri, M. Z. Pan, S. Yang, et al. (2025) Why do multi-agent LLM systems fail?. arXiv preprint arXiv:2503.13657. Cited by: 3rd item. S. Chen, Y. Li, Z. Xu, Y. Zeng, et al. (2025) DAST: context-aware compression in LLMs via dynamic allocation of soft tokens. In Findings of the Association for Computational Linguistics: ACL, Cited by: §3.3. F. Cuconasu, G. Trappolini, F. Siciliano, S. Filice, C. Campagnano, Y. Maarek, N. Tonellotto, and F. Silvestri (2024) The power of noise: redefining retrieval for RAG systems. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, Cited by: §2, §3.3. T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. RĂŠ (2022) FlashAttention: fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems, Vol. 35. Cited by: §2. T. Dao and A. Gu (2024) Transformers are SSMs: generalized models and efficient algorithms through structured state space duality. Proceedings of the International Conference on Machine Learning. Cited by: §6.1. Y. Ding, L. L. Zhang, C. Zhang, Y. Xu, N. Shang, J. Xu, F. Yang, and M. Yang (2024) LongRoPE: extending LLM context window beyond 2 million tokens. In Proceedings of the International Conference on Machine Learning, Note: arXiv:2402.13753 Cited by: §1, §2. Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, and H. Wang (2024) Retrieval-augmented generation for large language models: a survey. arXiv preprint arXiv:2312.10997. Cited by: §2. G. Gartner, D. A. Bennett, and T. Morita (2007) Towards ubiquitous cartography. Cartography and Geographic Information Science 34 (4), p. 247â257. Cited by: §1, §2. G. Gartner (2025) Why AI and large language models benefit from cartography. ArcNews. Note: Winter 2025 Cited by: §2. P. Glorioso, Q. Anthony, et al. (2024) The zamba2 suite: technical report. arXiv preprint arXiv:2411.15242. Cited by: §6.1. S. Goldstein and C. D. Kirk-Giannini (2024) A case for AI consciousness: language agents and global workspace theory. arXiv preprint arXiv:2410.11407. Cited by: §6.3. A. Gu and T. Dao (2024) Mamba: linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752. Cited by: §1, §6.1. S. Gu (2026) Long context, less focus: a scaling gap in LLMs revealed through privacy and personalization. arXiv preprint arXiv:2602.15028. Cited by: §3.1. M. Gupte, E. Dixit, M. Tayyab, and A. Adiththan (2025) What works for âlost-in-the-middleâ in LLMs? A study on GM-Extract and mitigations. arXiv preprint arXiv:2511.13900. Cited by: §2, 3rd item. X. Hou, Y. Zhao, et al. (2025) Model context protocol (MCP): landscape, security threats, and future research directions. arXiv preprint arXiv:2503.23278. Cited by: Table 4. C. Hsieh, Y. Chuang, C. Li, Z. Wang, L. Le, A. Kumar, J. Glass, A. Ratner, C. Lee, R. Krishna, and T. Pfister (2024) Found in the middle: calibrating positional attention bias improves long context utilization. In Findings of the Association for Computational Linguistics: ACL, Cited by: §2, §3.3. L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu (2023) A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232. Cited by: §3.1, Table 1. Z. Jing, Y. Su, Y. Han, et al. (2024) When large language models meet vector databases: a survey. arXiv preprint arXiv:2402.01763. Cited by: Table 1. S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, et al. (2022) Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221. Cited by: §3.1. D. Kahneman (2011) Thinking, fast and slow. Farrar, Straus and Giroux. Cited by: §6.2. M. Kang, W. Chen, D. Han, H. A. Inan, L. Wutschitz, Y. Chen, R. Sim, and S. Rajmohan (2025) ACON: optimizing context compression for long-horizon LLM agents. arXiv preprint arXiv:2510.00615. Cited by: §3.3, §3.5. O. Khattab, A. Singhvi, P. Maheshwari, Z. Zhang, K. Santhanam, S. Vardhamanan, S. Haq, A. Sharma, T. T. Joshi, H. Mober, et al. (2024) DSPy: compiling declarative language model calls into self-improving pipelines. In International Conference on Learning Representations, Note: arXiv:2310.03714 Cited by: Table 1. A. Kostka and J. A. Chudziak (2025) Towards cognitive synergy in LLM-based multi-agent systems: integrating theory of mind and critical evaluation. arXiv preprint arXiv:2507.21969. Cited by: 3rd item. Letta (2025) Introducing context repositories: git-based memory for coding agents. Note: https://w.letta.com/blog/context-repositoriesAccessed: 2026-02-27 Cited by: item ĎĎ (Simplification) = 4 (P,E,A,D), §3.5, §4.2. P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. KĂźttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela (2020) Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, Vol. 33. Cited by: §2. D. Li, S. Cao, T. Griggs, et al. (2025a) LLMs can easily learn to reason from demonstrations structure, not content, is what matters!. arXiv preprint arXiv:2502.07374. Cited by: §6.2. H. Li, Y. Li, A. Tian, T. Tang, et al. (2025b) A survey on large language model acceleration based on KV cache management. arXiv preprint arXiv:2412.19442. Cited by: §2, §3.2. M. Li, L.H. Xu, Q. Tan, L. Ma, T. Cao, and Y. Liu (2025c) Sculptor: empowering LLMs with cognitive agency via active context management. arXiv preprint arXiv:2508.04664. Cited by: §6.3. Z. Li et al. (2025) Large language models struggle to capture long-distance relational knowledge. In Findings of the Association for Computational Linguistics: NAACL, Cited by: §1, §3.2. Z. Li, S. Song, H. Wang, S. Niu, D. Chen, J. Yang, C. Xi, H. Lai, J. Zhao, Y. Wang, et al. (2025d) MemOS: an operating system for memory-augmented generation in large language models. arXiv preprint arXiv:2505.22101. Cited by: item Ď (Selection) = 5 (P,E,C,A,D), item Îą (Aggregation) = 4 (P,E,A,D), item Îť (Layering) = 5 (P,E,C,A,D), §2, §3.3, §4.3. Z. Li, C. Xi, C. Li, D. Chen, B. Chen, S. Song, S. Niu, H. Wang, J. Yang, C. Tang, et al. (2025e) MemOS: a memory OS for AI system. arXiv preprint arXiv:2507.03724. Cited by: item Ď (Selection) = 5 (P,E,C,A,D), item Ď (Projection) = 5 (P,E,C,A,D), §3.3, §4.3, 4th item. Z. Li, D. Zhang, M. Zhang, et al. (2025f) From system 1 to system 2: a survey of reasoning large language models. arXiv preprint arXiv:2502.17419. Cited by: §6.2. O. Lieber, B. Lenz, H. Bata, G. Cohen, J. Osin, I. Dalmedigos, E. Safahi, S. Meirom, Y. Belinkov, A. Shashua, and Y. Shoham (2024) Jamba: a hybrid transformer-mamba language model. arXiv preprint arXiv:2403.19887. Cited by: §6.1. H. Lin, M. Batty, S. E. Jørgensen, B. Fu, M. Konecny, A. Voinov, P. Torrens, G. Lu, A. Zhu, J. P. Wilson, et al. (2011) Cartography: challenges and potential in the virtual geographic environments era. Annals of GIS 17 (3), p. 135â148. Cited by: §2. T. Lindenbauer, I. Slinko, L. Felder, E. Bogomolov, and Y. Zharov (2025) The complexity trap: simple observation masking is as efficient as LLM summarization for agent context management. arXiv preprint arXiv:2508.21433. Cited by: 1st item, §3.3, §3.4. N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang (2024) Lost in the middle: how language models use long contexts. In Transactions of the Association for Computational Linguistics, Vol. 12, p. 157â173. Cited by: §1, §1, §2, §3.1, §3.2, §3.3, Table 1, §4.5, item 2. S. Liu et al. (2025) Memory in the age of AI agents: a survey. arXiv preprint arXiv:2512.13564. Cited by: §2, §3.1, §3.3. A. M. MacEachren (1995) How maps work: representation, visualization, and design. Guilford Press. Cited by: §1, §2, §3.2. R. B. McMaster and K. S. Shea (1992) Generalization in digital cartography. Association of American Geographers. Cited by: §1, §2, 2nd item, §3.3, §3.5. L. Mei, J. Yao, Y. Ge, Y. Wang, B. Bi, et al. (2025) A survey of context engineering for large language models. arXiv preprint arXiv:2507.13334. Cited by: §1, §2. N. Mu, J. Lu, et al. (2025) A closer look at system prompt robustness. arXiv preprint arXiv:2502.12197. Cited by: Table 4. S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J. Wen, and C. Li (2025) Large language diffusion models. arXiv preprint arXiv:2502.09992. Cited by: §1, §6.1. C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez (2023) MemGPT: towards LLMs as operating systems. arXiv preprint arXiv:2310.08560. Cited by: item ĎĎ (Simplification) = 4 (P,E,A,D), item Ď (Projection) = 5 (P,E,C,A,D), item Îť (Layering) = 5 (P,E,C,A,D), §2, §3.1, §4.2, 4th item. J. S. Park, J. C. OâBrien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein (2023) Generative agents: interactive simulacra of human behavior. In Proceedings of the ACM Symposium on User Interface Software and Technology, Cited by: §2. N. Paulsen (2025) Context is what you need: the maximum effective context window for real world limits of LLMs. arXiv preprint arXiv:2509.21361. Cited by: §1, §3.2, §5.1. B. Peng et al. (2025) RWKV-7 âgooseâ with expressive dynamic state evolution. arXiv preprint arXiv:2503.14456. Cited by: §1, §2, §6.1. M. Pink, Q. Wu, et al. (2025) Position: episodic memory is the missing piece for long-term LLM agents. arXiv preprint arXiv:2502.06975. Cited by: Table 4. A. H. Robinson and B. B. Petchenik (1976) The nature of maps: essays toward understanding maps and mapping. University of Chicago Press. Cited by: §1, §2. R. E. Roth, C. A. Brewer, and M. S. Stryker (2011) A typology of operators for maintaining legible map designs over multiple scales. Cartographic Perspectives (68), p. 29â64. Cited by: §2, 1st item, §3.3, §3.5. T. Schick, J. Dwivedi-Yu, R. DessĂŹ, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom (2023) Toolformer: language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, Vol. 36. Cited by: §2. C. Shao, F. Meng, and J. Zhou (2025) Beyond next token prediction: patch-level training for large language models. In International Conference on Learning Representations, Cited by: §6.2. F. Shi, X. Chen, K. Misra, N. Scales, D. Dohan, E. Chi, N. Schärli, and D. Zhou (2023) Large language models can be easily distracted by irrelevant context. Proceedings of the International Conference on Machine Learning. Cited by: §2, §3.3. J. P. Snyder (1987) Map projections: a working manual. U.S. Geological Survey Professional Paper 1395, U.S. Government Printing Office. Cited by: item 2, §3.2, §3.4. J. Stoter, M. Post, V. van Altena, M. Rumor, and C. Rizos (2014) State-of-the-art of automated generalisation in commercial software. Cartographica 49 (1), p. 60â71. Cited by: §2. T. R. Sumers, S. Yao, K. Narasimhan, and T. L. Griffiths (2023) Cognitive architectures for language agents. arXiv preprint arXiv:2309.02427. Cited by: Table 4. J. Sweller (1988) Cognitive load during problem solving: effects on learning. Cognitive Science 12 (2), p. 257â285. Cited by: §3.1. K. Tran, D. Dao, M. Nguyen, Q. Pham, B. OâSullivan, and H. D. Nguyen (2025) Multi-agent collaboration mechanisms: a survey of LLMs. arXiv preprint arXiv:2501.06322. Cited by: §6.3. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ĺ. Kaiser, and I. Polosukhin (2017) Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30. Cited by: Definition 3. P. Vasylenko, H. Pitorro, A. F. T. Martins, and M. Treviso (2025) Long-context generalization with sparse attention. arXiv preprint arXiv:2506.16640. Cited by: §3.2. Volcengine (2025) OpenViking: an open-source context database for AI agents. Note: https://github.com/volcengine/OpenVikingAccessed: 2026-02-27 Cited by: item Ď (Selection) = 5 (P,E,C,A,D), item Ď (Projection) = 5 (P,E,C,A,D), item Îť (Layering) = 5 (P,E,C,A,D), §3.3, §3.5, §4.4. C. von Clausewitz (1832) On war. DĂźmmlers Verlag. Note: Translated by Michael Howard and Peter Paret, Princeton University Press, 1976 Cited by: §1. E. Wallace, K. Xiao, R. Leike, et al. (2024) The instruction hierarchy: training LLMs to prioritize privileged instructions. arXiv preprint arXiv:2404.13208. Cited by: §3.3. G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar (2023a) Voyager: an open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291. Cited by: Table 1. K. Wang, Y. Lin, J. Lou, et al. (2026) E-mem: multi-agent based episodic context reconstruction for LLM agent memory. arXiv preprint arXiv:2601.21714. Cited by: §3.3. Q. Wang, Y. Fu, et al. (2023b) Recursively summarizing enables long-term dialogue memory in large language models. arXiv preprint arXiv:2308.15022. Cited by: Table 4. W. Wang et al. (2026) Intelligence degradation in long-context LLMs: critical threshold determination via natural length distribution analysis. arXiv preprint arXiv:2601.15300. Cited by: §3.2. J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou (2022) Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, Vol. 35. Cited by: §2. X. Wu, Y. Wang, S. Jegelka, and A. Jadbabaie (2025) On the emergence of position bias in transformers. In Proceedings of the International Conference on Machine Learning, Cited by: §2, §3.2, §3.3. Y. Wu, M. N. Rabe, D. Hutchins, and C. Szegedy (2022) Memorizing transformers. In International Conference on Learning Representations, Cited by: §2. G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis (2024) Efficient streaming language models with attention sinks. In International Conference on Learning Representations, Cited by: §3.2. Y. Xiao et al. (2025) Improving the efficiency of LLM agent systems through trajectory reduction. arXiv preprint arXiv:2509.23586. Cited by: §3.3. Z. Xiong, Y. Lin, et al. (2025) How memory management impacts LLM agents: an empirical study of experience-following behavior. arXiv preprint arXiv:2505.16067. Cited by: Table 1. W. Xu, Z. Liang, K. Mei, H. Gao, J. Tan, and Y. Zhang (2025) A-MEM: agentic memory for LLM agents. In Advances in Neural Information Processing Systems, Cited by: §3.3. C. Yang et al. (2025) Training LLMs beyond next token prediction â filling the mutual information gap. arXiv preprint arXiv:2511.00198. Cited by: §6.2. S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao (2023) ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations, Cited by: §2, Table 1. J. Ye, Z. Xie, L. Zheng, J. Gao, Z. Wu, X. Jiang, Z. Li, and L. Kong (2025) Dream 7b: diffusion large language models. arXiv preprint arXiv:2508.15487. Cited by: §6.1. Y. Yu, H. Jiang, X. Luo, Q. Wu, C. Lin, D. Li, Y. Yang, Y. Huang, and L. Qiu (2025) Mitigate position bias in large language models via scaling a single dimension. In Findings of the Association for Computational Linguistics: ACL, Cited by: §2. S. Yuen, F. Gomez Medina, T. Su, et al. (2025) Intrinsic memory agents: heterogeneous multi-agent LLM systems through structured contextual memory. arXiv preprint arXiv:2508.08997. Cited by: §6.3. A. L. Zhang, T. Kraska, and O. Khattab (2025a) Recursive language models. arXiv preprint arXiv:2512.24601. Cited by: §6.1. Q. Zhang, C. Hu, S. Upasani, B. Ma, et al. (2025b) Agentic context engineering: evolving contexts for self-improving language models. arXiv preprint arXiv:2510.04618. Cited by: 2nd item, 1st item, §3.1, §4.5, item 4. W. Zhang, L. Zeng, et al. (2025c) AgentOrchestra: orchestrating multi-agent intelligence with the tool-environment-agent (TEA) protocol. arXiv preprint arXiv:2506.12508. Cited by: Table 4. Z. Zhang, R. Chen, S. Liu, Z. Yao, O. Ruwase, B. Chen, X. Wu, and Z. Wang (2024) Found in the middle: how language models use long contexts better via plug-and-play positional encoding. In Advances in Neural Information Processing Systems, Note: arXiv:2403.04797 Cited by: §2. Y. Zhou, Y. Lei, S. Si, et al. (2025) From context to EDUs: faithful and structured context compression via elementary discourse unit decomposition. arXiv preprint arXiv:2512.14244. Cited by: 3rd item, §3.3, §3.5. A. Zweiger, J. Pari, H. Guo, E. AkyĂźrek, Y. Kim, and P. Agrawal (2025) Self-adapting language models. In Advances in Neural Information Processing Systems, Cited by: §6.1. Appendix A Operator Implementation Scoring Evidence This appendix provides the evidence basis for each cell in TableË5. Each systemâoperator pair is evaluated on five binary criteria: Present (any mechanism performs this transformation), Explicit (named architectural component), Configurable (parameters tunable by user/developer), Automated (occurs without manual intervention), and Documented (described in official documentation or papers). Scores range from 0 (absent) to 5 (all criteria met). A.1 Claude Code Ď (Reconnaissance) = 5 (P,E,C,A,D) Subagent spawning via the Task tool is the core exploration mechanism [Anthropic, 2025]. P: Subagents traverse âŹB by executing file searches, code reads, and web fetches in isolated contexts. E: Named as âTask toolâ with typed agent variants (Explore, Plan, general-purpose). C: Users specify agent type, prompt, and isolation mode (e.g., worktree). A: The orchestrator autonomously decides when to spawn subagents. D: Documented in Claude Code official documentation. Ď (Selection) = 2 (P,A) P: Subagents implicitly filter what to return to the orchestratorânot all explored content is surfaced. A: This filtering is automated within the subagentâs reasoning. Not E: no named âselectionâ component exists. Not C: no user-facing parameter controls filtering granularity. Not D: not described as a distinct feature. ĎĎ (Simplification) = 4 (P,E,A,D) P: Subagent results are compressed before return to the orchestrator. E: Described as âcondensed resultsâ in documentation. A: Compression is automaticâsubagents return summaries, not raw tool output. D: Documented as part of the subagent architecture. Not C: no user-configurable compression level or summary length parameter. Îą (Aggregation) = 1 (P) P: When multiple subagents return results, their outputs coexist in the orchestratorâs context, but there is no explicit mechanism to fuse or deduplicate them. Not E,C,A,D. Ď (Projection) = 2 (P,A) P: Subagent results are projected from the subagentâs internal representation into the orchestratorâs token sequence. A: This happens automatically upon subagent completion. Not E: no named projection schema. Not C: results are returned at a single fixed resolutionâno multi-resolution control. Not D: not described as a distinct projection mechanism. δ (Displacement) = 2 (P,A) P: Subagent summaries enter the orchestratorâs context at the recency-biased end of the sequence, implicitly placing them in a high-salience position. A: This is an automatic consequence of append-only context. Not E: no named reordering mechanism. Not C: position is determined by insertion order, not configurable. Not D: not described as positional management. Îť (Layering) = 5 (P,E,C,A,D) P: System prompts, user messages, tool results, and CLAUDE.md instructions occupy distinct typed namespaces. E: Named layers include âsystem-reminder,â âtool results,â and âuser messages.â C: Users configure layers via CLAUDE.md, system prompts, and hook configurations. A: Layer separation is enforced automatically by the runtime. D: Documented in Claude Code documentation. A.2 Letta (MemGPT) Ď (Reconnaissance) = 1 (P) P: Letta agents can invoke tools and external APIs. Not E,C,A,D: tool use is a general capability, not an architecturally distinct exploration mechanism. Ď (Selection) = 2 (P,A) P: Recall memory search retrieves relevant archived content. A: Search is triggered automatically during context assembly. Not E: selection is embedded within the memory hierarchy, not a standalone component. Not C,D. ĎĎ (Simplification) = 4 (P,E,A,D) P: Evicted content undergoes recursive summarization [Packer et al., 2023]. E: Named as ârecursive summarizationâ in the MemGPT paper. A: Triggered automatically when context overflows. D: Described in the MemGPT paper and Letta documentation [Letta, 2025]. Not C: summarization parameters are not user-configurable. Îą (Aggregation) = 2 (P,A) P: Sleep-time background processes reorganize and merge related memory entries. A: Runs asynchronously without user intervention. Not E: not a named standalone component. Not C,D. Ď (Projection) = 5 (P,E,C,A,D) P: Core/recall memory tiers implement multi-resolution projection [Packer et al., 2023]. E: Named as âcore memoryâ (full resolution) and ârecall memoryâ (compressed resolution). C: Users configure core memory block contents and structure. A: Projection from recall to context is automatic. D: Central architectural contribution of the MemGPT paper. δ (Displacement) = 1 (P) P: Content position is determined by insertion order and memory tier, providing implicit positional bias. Not E,C,A,D: no explicit positional management mechanism. Îť (Layering) = 5 (P,E,C,A,D) P: Core memory blocks (human, persona, system) are typed namespaces [Packer et al., 2023]. E: Named block types with distinct roles. C: Users define and modify block contents. A: Block separation is maintained automatically. D: Documented in MemGPT paper and Letta platform. A.3 MemOS Ď (Reconnaissance) = 1 (P) P: Agents can query external sources. Not E,C,A,D: general capability, not an architecturally distinct exploration mechanism. Ď (Selection) = 5 (P,E,C,A,D) P: DAG-based scheduling selects which memory nodes to activate [Li et al., 2025d, e]. E: Named as âMemSchedulerâ with DAG-based dependency resolution. C: Users configure scheduling policies and memory access patterns. A: Next-scene prediction preloads memory before explicit request. D: Described in both MemOS papers. ĎĎ (Simplification) = 2 (P,A) P: Graph operations may reduce information during traversal. A: Occurs during automated graph queries. Not E: simplification is not a named standalone mechanism. Not C,D. Îą (Aggregation) = 4 (P,E,A,D) P: Graph-structured memory aggregates related entities into composite nodes [Li et al., 2025d]. E: MemCube merges are explicit graph operations. A: Aggregation occurs during graph maintenance. D: Described in MemOS papers. Not C: aggregation policies are not directly user-configurable. Ď (Projection) = 5 (P,E,C,A,D) P: Graph â text projection with multiple schemas [Li et al., 2025e]. E: Named projection modes include entity lists, relationship summaries, and subgraph traversals. C: Projection schema is selectable per query. A: Projection is automatic during memory retrieval. D: Described in both MemOS papers. δ (Displacement) = 1 (P) P: No explicit positional management; content order is determined by graph traversal order. Not E,C,A,D. Îť (Layering) = 5 (P,E,C,A,D) P: MemCube abstraction isolates knowledge bases as typed containers [Li et al., 2025d]. E: Named as âMemCubeâ with composable isolation. C: Users create and configure distinct MemCubes. A: Isolation is maintained automatically. D: Central contribution of the MemOS architecture. A.4 OpenViking Ď (Reconnaissance) = 1 (P) P: Agents can traverse the filesystem hierarchy to discover content. Not E,C,A,D: traversal is a general retrieval mechanism, not a named exploration component. Ď (Selection) = 5 (P,E,C,A,D) P: Path-based directory traversal selects content by hierarchical location [Volcengine, 2025]. E: Named as structured traversal with custom URI scheme. C: Users configure directory structure and retrieval paths. A: Selection follows directory recursion automatically. D: Documented in OpenViking repository. ĎĎ (Simplification) = 1 (P) P: L0 summaries (âź 100 tokens) are brief, but this is a property of the projection tier, not a standalone simplification mechanism. Not E,C,A,D as a distinct operator. Îą (Aggregation) = 1 (P) P: No explicit aggregation mechanism; retrieved content is presented individually. Not E,C,A,D. Ď (Projection) = 5 (P,E,C,A,D) P: L0/L1/L2 tiered loading implements multi-resolution projection [Volcengine, 2025]. E: Named tiers with explicit token budgets (L0: âź 100, L1: âź 1,000, L2: full). C: Tier selection is configurable per retrieval request. A: Tier assignment is automatic based on query context. D: Documented as the core architecture. δ (Displacement) = 1 (P) P: No explicit positional management; retrieval order follows filesystem hierarchy. Not E,C,A,D. Îť (Layering) = 5 (P,E,C,A,D) P: Hierarchical filesystem with custom URI scheme enforces namespace separation [Volcengine, 2025]. E: Named directory hierarchy with typed content. C: Users configure the filesystem structure. A: Namespace isolation is maintained by the filesystem abstraction. D: Documented in OpenViking repository. Appendix B Lean 4 Formalization Five structural properties of the framework have been formalized and mechanically verified in Lean 4 using Mathlib. All proofs compile with only standard axioms (propext, Classical.choice, Quot.sound). Theorem statements are shown below; full proofs are available at https://github.com/lucifer1004/context-cartography. B.1 Unique Zone Membership (Definition 1) Every element belongs to exactly one zone, derived from the disjointness and coverage axioms of the ContextState structure. ⏠theorem mem_unique_zone (s : ContextState U) (x : U) : (x â s.blackFog â§ x â s.grayFog â§ x â s.visible) ⨠(x â s.blackFog â§ x â s.grayFog â§ x â s.visible) ⨠(x â s.blackFog â§ x â s.grayFog â§ x â s.visible) B.2 Boundary Operator Uniqueness (§3.5) Each boundary transition is governed by exactly one operator kind. Removing any boundary operator leaves its transition without representational governance. ⏠-- Ď is the sole operator for B â G theorem reconnaissance_sole : â op, operatorCovers op (blackFog, grayFog) â op = .reconnaissance -- Ď is the sole operator kind for G â B theorem selection_sole : â op, operatorCovers op (grayFog, blackFog) â â m, op = .selection m -- Ď+Ď^+ is the sole non-selection operator for G â V theorem forwardProjection_sole : â op, operatorCovers op (grayFog, visible) â op = .forwardProjection ⨠â m, op = .selection m -- ĎâĎ^- is the sole non-selection operator for V â G theorem inverseProjection_sole : â op, operatorCovers op (visible, grayFog) â op = .inverseProjection ⨠â m, op = .selection m B.3 Composition Non-Commutativity (§3.5) A concrete witness over Fin 3 proving that δâĎâ Ďâδ Ďâ Ď Î´: ĎĎ keeps only element 0, δ adds element 1. Then δâ(Ďâ(0,2))=0,1δ(Ď(\0,2\))=\0,1\ but Ďâ(δâ(0,2))=0Ď(δ(\0,2\))=\0\. ⏠theorem composition_noncommutative : δ1 _1(Ď0 _0.simplify 0, 2) â Ď0 _0.simplify (δ1 _1 0, 2) B.4 Context Collapse (§3.1) Destructive compaction (without archival) provably loses information to âŹB; archival compaction preserves all content in âŞG . ⏠-- Destructive: non-summary V content moves to B theorem destructive_loses_to_blackfog (s : ContextState U) (summary : Set U) (x : U) (hx_vis : x â s.visible) (hx_not : x â summary) : x â (destructiveCompaction s summary).newBlackFog -- Archival: all original V content remains in G ⪠V theorem archival_preserves_all (s : ContextState U) (summary : Set U) : s.visible â (archivalCompaction s summary).newGrayFog ⪠(archivalCompaction s summary).newVisible B.5 Invariant Impossibility (§3.4) No strict reduction can simultaneously preserve semantic content (critical elements retained), structural integrity (linked elements co-occur), and salience priority (high-priority element retained). Over Fin 3, element 0 is semantically critical, element 0 structurally requires element 1, and element 2 is high-priority. Any strict reduction from 0,1,2\0,1,2\ must drop at least one, violating at least one invariant. ⏠theorem invariant_impossibility : ÂŹ â (f : Set (Fin 3) â Set (Fin 3)), f 0, 1, 2 â 0, 1, 2 â§ preservesSemantic f (â ¡ = 0) 0, 1, 2 â§ preservesStructural f (fun x y â x = 0 â§ y = 1) 0, 1, 2 â§ preservesSalience f 2 0, 1, 2 The proof proceeds by deriving 0âfâ(S)0â f(S) from semantic preservation, 1âfâ(S)1â f(S) from structural preservation, and 2âfâ(S)2â f(S) from salience preservation, yielding fâ(S)âSf(S) Sâcontradicting strict reduction.