Paper deep dive
Toward Personal Intelligence Through Cooperative Observation
Yashar Talebirad, Osman Jime, Ali Parsaee, Eden Redman, Yongbin Kim, Osmar R. Zaiane
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/19/2026, 4:16:40 AM
Summary
The paper introduces 'cooperative observation' as a framework for personal intelligence, where a personal AI system and its user jointly govern the system's access to user context. This feedback loop involves the system building a model of the user, the user evaluating actions, and the user controlling future observation channels based on usefulness, trust, and privacy. The authors propose an information-theoretic bound on personalization quality based on observation history and describe a prototype system called Organizm.
Entities (7)
Relation Signals (6)
Cooperative Observation → defines → Personal Intelligence
confidence 95% · We use the term cooperative observation for this feedback loop... and propose it as a framework for personal intelligence.
Cooperative Observation → involves → User Control
confidence 92% · the user’s consent and control shape what it can observe next.
Cooperative Observation → involves → User Evaluation
confidence 92% · the user evaluates its actions
Personal AI → requires → User Model
confidence 90% · A personal AI system needs a model of the user's goals, constraints, and ongoing commitments
Observation History → bounds → Personalization Quality
confidence 88% · the quality of that model is bounded by what the system can observe.
Organizm → implements → Cooperative Observation
confidence 85% · We report a preliminary single-subject account from Organizm, a prototype used over six months
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:A personal AI system needs a model of the user's goals, constraints, and ongoing commitments to plan and act on their behalf, and the quality of that model is bounded by what the system can observe. Broader observation does not by itself improve assistance because a bounded system must select and compress information for the task at hand. We argue that this observation bottleneck has a cooperative structure: the system builds a partial model of the user's changing life, the user evaluates its actions, and the user's consent and control shape what it can observe next. Useful and inspectable behavior can give users a reason to maintain or expand the observation channel, while failures can lead them to correct, narrow, revoke, or abandon it. We use the term cooperative observation for this feedback loop among usefulness, trust, and future access, and propose it as a framework for personal intelligence. We report a preliminary single-subject account from Organizm, a prototype used over six months, and outline evaluation directions for measuring how observation quality shapes personal AI.
Tags
Links
- Source: https://arxiv.org/abs/2608.17128v1
- Canonical: https://arxiv.org/abs/2608.17128v1
Trouble viewing inline? Open PDF directly →
Full Text
61,984 characters extracted from source content.
Expand or collapse full text
PILA ’26 (KDD 2026), August 9, 2026, Jeju Island, Republic of Korea Toward Personal Intelligence Through Cooperative Observation Yashar Talebirad Alberta Machine Intelligence Institute, University of Alberta Edmonton, Canada talebira@ualberta.ca Osman Jime MacEwan University Edmonton, Canada jimeo@mymacewan.ca Ali Parsaee University of Alberta Edmonton, Canada parsaee@ualberta.ca Eden Redman Network for Applied Technology Edmonton, Canada eden@nat.ltd Yongbin Kim University of Alberta Edmonton, Canada yongbin2@ualberta.ca Osmar R. Zaïane Alberta Machine Intelligence Institute, University of Alberta Edmonton, Canada zaiane@ualberta.ca Abstract A personal AI system needs a model of the user’s goals, constraints, and ongoing commitments to plan and act on their behalf, and the quality of that model is bounded by what the system can observe. Broader observation does not by itself improve assistance because a bounded system must select and compress information for the task at hand. We argue that this observation bottleneck has a cooperative structure: the system builds a partial model of the user’s changing life, the user evaluates its actions, and the user’s consent and control shape what it can observe next. Useful and inspectable behavior can give users a reason to maintain or expand the observation channel, while failures can lead them to correct, narrow, revoke, or abandon it. We use the term cooperative observation for this feedback loop among usefulness, trust, and future access, and propose it as a framework for personal intelligence. We report a preliminary single- subject account from Organizm, a prototype used over six months, and outline evaluation directions for measuring how observation quality shapes personal AI. Keywords personal intelligence, personal AI, cooperative observation, user modeling, privacy, agentic AI 1 Introduction Personal AI is moving from a speculative idea to an active engineer- ing target. Open-source projects such as OpenClaw and NanoClaw 1 aim to give users personal assistants that run across devices and tools, often emphasizing local control and inspectability. Commer- cial systems are also converging on persistent user memory, voice interfaces, and agentic tool use. At the same time, consumer hard- ware is moving closer to continuous capture through smart glasses, wearable audio, health sensors, and eventually neural interfaces. The need for user context links these developments. Personal intelligence depends on knowledge of the user’s goals, constraints, habits, and ongoing commitments. Relevant context lets a model plan, catch inconsistencies, and support decisions across time. The system cannot directly access the user’s life, so its personal model is bounded by a lossy stream of reports, sensor readings, and in- teractions. Agent software and sensing hardware are approaching 1 OpenClaw: https://github.com/openclaw/openclaw. NanoClaw: https://github.com/ nanocoai/nanoclaw. this observation bottleneck from opposite sides: agents need richer context to act personally, while new devices capture context that becomes useful once an agent can act on it. Much of the relevant context is already being collected: smart watches and health platforms record sleep, activity, and stress, sig- nals directly useful for personal assistance, but access is often re- stricted by platform-specific exports, APIs, and permission systems. An assistant chosen by the user therefore lacks access to signals that would help it most, and the user may have no practical way to transfer data describing their own body and days. Richer sensing devices expand what can enter the observation stream, but leave open how that stream’s configuration is governed and which objec- tive is optimized over the resulting history. Local processing can support cooperative observation by keeping channel configuration and memory under the user’s control, allowing continuous capture without continuous external disclosure. Computational systems already model people at scale. Recom- mender systems, social feeds, and advertising infrastructures infer preferences and vulnerabilities from behavioral traces, and the ob- jectives they serve, such as engagement, retention, and revenue, belong to the platform rather than to the person being modeled [2,32–34]. Richer observation flowing into systems of this kind would deepen surveillance, manipulation, and dependency. The same modeling capacity, aimed at goals the user sets and evaluated under the user’s own reward signal, could instead help a person understand and steer their own life amid competing models built for other objectives, if the observation channel and the objective remain under that person’s control. Personal intelligence depends on structural conditions that make observation cooperative. In a partially observable decision problem, an assistant needs three components: a model of the relevant state, a signal that defines whether its actions are good, and an observation process that deter- mines what it can learn next. In personal AI, all three are anchored in the same person: the state concerns the user’s goals, constraints, commitments, and other parts of life relevant to the task; the user’s own evaluation supplies feedback; and their willingness to share or withhold data governs much of the observation process. Most personalization systems split these roles across parties: platforms often define the objective and govern how behavioral data is used. In user-owned personal AI, however, the person being modeled can define success and control future access to their data. 1 arXiv:2608.17128v1 [cs.AI] 17 Aug 2026 PILA ’26 (KDD 2026), August 9, 2026, Jeju Island, Republic of KoreaYashar Talebirad, Osman Jime, Ali Parsaee, Eden Redman, Yongbin Kim, and Osmar R. Zaïane We propose cooperative observation as a framework for user- governed personal intelligence. Evaluation therefore considers both the quality of assistance and changes to the observation channel. We use an observation-channel spectrum to illustrate how available signals may broaden from deliberate reports to continuous and neu- ral observation. We also frame task-conditioned context allocation as the problem of selecting and representing user information for a decision under processing and disclosure constraints. We then describe Organizm, a prototype built on user-owned files, explicit memory, and feedback-driven planning. 2 The Rise of Personal AI Systems Large language models can act as general-purpose interfaces over tools, files, and communication channels. Reinforcement learning from human feedback also shows that model behavior can change substantially when training is shaped by human preferences and instructions [5,25]. These capabilities support assistants that plan, maintain memory, incorporate feedback, and use tools across sus- tained interactions. Recent agent systems combine language models with mem- ory, planning, and reflection, showing how persistent context can change the behavior of language-model systems [27]. Projects such as OpenClaw and NanoClaw extend this in a user-facing direction, offering cross-platform personal agents with persistent memory, inspectable execution, and local or user-chosen models. OpenClaw- style ecosystems already share reusable skills that extend what an agent can do. 2 Observation and memory can be modular as well, with each module exposing what it collects and retains. The con- text a personal assistant needs most, covering goals, habits, health, and finances, is also the context users are least willing to place in services they cannot inspect. Model capability is also becoming more deployment-efficient. An analysis of 51 open-source base models found that maximum capa- bility per parameter increased sharply over time, while architecture- aware scaling experiments show that models can gain accuracy and inference throughput under matched training budgets [1,43]. As this frontier advances, personal agents can interpret and act on a fixed history more effectively while keeping more inference on user-controlled devices. This complements observation-channel growth, which can supply additional user-specific evidence. A multi-week field deployment of a context-aware LLM chat- bot found that participants who communicated more frequently about their context received more specific and relevant sugges- tions. Inaccurate recommendations discouraged some users from further interaction, while privacy concerns and reluctance to criti- cize the chatbot also limited context sharing and correction [44]. The support that the users wanted changed during the deployment, moving from action discovery toward planning, tracking, reflection, motivation, and accountability. Devices are also becoming better at capturing the user’s world. Smart glasses, wearable audio, and health sensors already expose richer context than typed text, and neural interfaces have begun to decode attempted speech directly from neural activity [42]. These devices widen the observation channel available to personal AI. 2 OpenClaw skills documentation: https://docs.openclaw.ai/tools/skills (accessed July 2026). Personal intelligence also requires continual learning. The user’s state is non-stationary as goals change, preferences drift, and con- straints appear. Continual-learning research studies how a system can absorb new information without overwriting what it already knows [26], and a personal AI faces that trade-off in its user model. Persistent memory and feedback incorporation help distinguish stable preferences from temporary noise. The same dependence on context also operates within a single user-system relationship. A user who begins by sharing only explicit tasks may later share priorities, constraints, or project history after receiving useful plans. The user may connect calendars, wearables, or continuous capture as trust grows, then narrow or close the channel after failures. The framework treats user control over these changes as a requirement. 3 The Cooperative Observation Framework A personal AI acts under partial observability [17] and under uncer- tainty about what its user values [12]. It must infer the task-relevant parts of the user’s changing life while learning their goals and pref- erences from feedback. Standard partially observable models specify the observation process in advance, while systems with controllable sensing may allow the agent to choose what to observe. In user-owned personal AI, much of this choice belongs to the user, who can enable, with- hold, or revoke data sources in response to system performance. The same person is therefore the subject of the state model, the evaluator of actions, and the controller of the observation chan- nel. In the formal setup below, uppercase letters denote random variables and lowercase letters denote realized values. The framework has three parts. A one-step decision model spec- ifies the user’s state, the system’s observations and actions, and the user’s immediate and reflective evaluations. An information- theoretic bound then describes the ceiling imposed by the system’s observation history. Finally, a feedback loop lets the user’s experi- ence of the action change what the system may observe next. 3.1 A One-Step Decision Model State. We use푆 푡 ∈ Sfor the user’s evolving state: goals, con- straints, commitments, ongoing priorities, external facts such as deadlines and meetings, and internal variables such as motivation, overload, and preference drift. A particular task or outcome deter- mines which components of this state are relevant to choosing or evaluating an action: a calendar may reveal a scheduling conflict while carrying little information about motivation. This creates an information-bottleneck problem: a bounded system must pre- serve the information needed for the decision while compressing a broader observation history [38]. Section 4.2 develops this connec- tion. Internal variables are latent and must be inferred from reports or behavior, whereas external facts become available through con- nected data sources or user reports. Observation. At each step, the system receives an observation푂 푡 about푆 푡 . The channel configurationΦ 푡 specifies which data sources are active and how much the user engages with each, and it can change over time under the user’s control. For simplicity, we omit additional dependencies on prior actions and interaction history. 2 Toward Personal Intelligence Through Cooperative ObservationPILA ’26 (KDD 2026), August 9, 2026, Jeju Island, Republic of Korea Action. Let퐻 푡 = (푂 1:푡 ,Φ 1:푡 )denote the system’s information history, comprising the observations it has received and the channel configurations through which they arrived. Requests and queries are themselves observations, so퐻 푡 carries what the system has been asked as well as what it has learned, and derived memory such as summaries and indexes is a function of the same history. The system selects actions according to a policy휋, with퐴 푡 = 휋(퐻 푡 ). The policy may encode arbitrary general knowledge, but its user- specific information can come only from 퐻 푡 . Reward. The user evaluates the action through ratings, accep- tance decisions, or corrections. We write푅 푡 ∈Rfor a scalar sum- mary of the immediate evaluation available to the system. Cor- rections can also supply information for subsequent decisions. In this framework,푅 푡 represents user-provided evaluation rather than platform-specific objectives such as engagement, retention, or click- through rate. Preference-learning systems show one route for in- corporating such feedback into model behavior [5, 25]. Approval and reflective benefit. Immediate approval can differ from reflective assessment. We write푈 푡 ∈Rfor a scalar summary of the user’s reflective assessment of whether the action was bene- ficial, evaluated at a specified horizon. Unlike푅 푡 ,푈 푡 is not assumed to be available to the system during interaction because relevant consequences may become apparent only later. When푅 푡 is opti- mized as a proxy for푈 푡 , this becomes an instance of Goodhart’s law: 푅 푡 can rise without a corresponding gain in reflective benefit [11]. Section 6 outlines possible operational proxies for these quantities. 3.2 The Observation Ceiling For fixed model parameters and policy, the system acts on its in- formation history, and the information flow forms a Markov chain 푆 푡 → 퐻 푡 → 퐴 푡 . By the data processing inequality [8]: 퐼(푆 푡 ; 퐴 푡 ) ≤ 퐼(푆 푡 ; 퐻 푡 ).(1) The left side measures how much the system’s action varies with the user’s true state under the chosen state distribution. An action independent of the state carries zero mutual information, while greater state dependence can produce higher mutual information. Mutual information alone does not imply utility, since an action can encode state without helping the user. The right side measures how much the information history reveals about the current state. The inequality says that state-dependent personalization cannot exceed the information the history carries about the user. The inequality gives only an upper bound: showing that observation limits a particular system requires comparing the same system under different observation conditions. The bound separates two sources of improvement. At a given step, changing the model or policy can help the system use in- formation already present in퐻 푡 , but it cannot add user-specific information to that history. Expanding the observation channel can raise퐼(푆 푡 ;퐻 푡 )and therefore the upper bound on state-dependent personalization. Which parts of the history are informative also de- pends on the task and time: old observations may lose relevance as the user’s state shifts, and different channels reveal different parts of the user’s life. The observation-channel ablation in Section 6 therefore holds the model fixed while varying the channel. User observation System memory/model System action User evaluation 푂 푡 퐻 푡 퐴 푡 푅 푡 Value, trust, effort, privacy→Φ 푡+1 Figure 1: Cooperative observation as a recurring loop. User evaluation affects the next observation channel, so observ- ability is partly earned through usefulness, trust, and control. 3.3 The Cooperative Feedback Loop The user’s experience can then change future observability. After seeing the system’s action, the user may maintain, expand, nar- row, or revoke the channel configuration for the next step,Φ 푡+1 . This decision depends on immediate evaluation, expectations about longer-term benefit, trust, effort, and privacy cost, and we leave the update rule unspecified. These factors can conflict: a user may find a system useful and still refuse to share health data when the privacy cost outweighs the marginal benefit. Channel configuration affects what the system can learn and how well it can help. Performance then affects the user’s willingness to maintain, expand, narrow, or revoke the channel. The user’s benefit also depends on who sets the objective and governs the channel. For instance, a richer user model can help a student notice weeks of drift from stated goals. Under a platform objective, the same model can target advertising at that drift and deepen dependence. At the level of individual updates, we define a channel update as cooperative when it remains under user control and improves reflective benefit. In contrast, expansion is extractive when access or engagement increases without such improvement. Continued use or disclosure alone does not establish trust or reflective benefit. We define cooperative observation as a feedback process in which a personal AI system and its user jointly adjust the system’s access to task-relevant user context because doing so benefits the user. Fig- ure 1 summarizes the process: the user provides a task, correction, or sensor stream; the system updates its working model and acts; and the user evaluates the action before sharing more, correcting the model, or withholding data. Cooperative observation has a cold-start problem: personal con- text can make the system more useful, but users may wait for evidence of usefulness and trustworthiness before sharing it. In a multi-week deployment of a context-aware chatbot, participants who communicated more context received more specific and rel- evant suggestions, while inaccurate suggestions and privacy con- cerns discouraged further interaction or disclosure [44]. The GOD 3 PILA ’26 (KDD 2026), August 9, 2026, Jeju Island, Republic of KoreaYashar Talebirad, Osman Jime, Ali Parsaee, Eden Redman, Yongbin Kim, and Osmar R. Zaïane model addresses the same problem by rewarding users with plat- form tokens for data that improves their assistant [29]. Cooperative observation changes this incentive loop by keeping evaluation and access decisions with the person being modeled. Useful assistance may itself motivate the user to maintain or expand the channel when the expected value of an added source for a particular task outweighs its effort and privacy costs. Viewed as a dynamical sys- tem, the loop may settle into a stable set of channels or remain history-dependent, for example when repeated success is needed to earn access but one serious failure is enough to lose it. Charac- terizing these regimes, including revocation and abandonment, is an open empirical question. 4 Observation Channels and the Complementary Self-Model Observation fidelity concerns how broadly, continuously, and di- rectly information about the user’s evolving state reaches the sys- tem. It can increase through more frequent observations, additional modalities, or more direct measurements. A continuous audio- visual stream may therefore raise the information ceiling beyond sporadic reports. Yet additional bandwidth need not carry informa- tion relevant to a particular task, so observation fidelity alone does not imply usefulness. Section 4.2 explains how a bounded system selects and compresses observations for a particular decision. We use four phases to mark broad positions on this spectrum, from deliberate self-report to continuous and neural observation (Figure 2). The phases are not mutually exclusive categories: a system may combine channels from several phases. Moving toward later phases can increase the information available to the system, while also raising privacy, consent, and governance demands. Phase 1: Manual reports. At the most deliberate end of the spec- trum, the user writes or says what matters: tasks, goals, and re- flections. This channel directly captures stated intentions, but it is low bandwidth, effortful, and prone to reporting gaps. With con- sistent reporting, it supports planning, goal tracking, and memory. Personal informatics systems provide a familiar example [22]. Phase 2: Automated behavioral context. The system observes cal- endars, messages, health data, or screen activity. This phase reduces reporting burden and provides evidence about what the user ac- tually did, supporting habit detection and behavior-intention gap analysis. It also increases privacy risk because passive traces reveal more than the user may consciously intend to share. Phase 3: Continuous multimodal capture. Smart glasses, always- on audio, cameras, and richer wearables capture situations the user may not explicitly report. Earlier lifelogging and wearable-camera systems, including MyLifeBits and SenseCam, explored related ideas of personal capture and retrospective memory support [9,14]. This phase can support context-aware assistance and fine-grained situation recognition. However, continuous capture also collects information about people other than the user, making bystander pri- vacy, consent, and transparency about inferred information central design requirements. Phase 4: Neural or near-neural interfaces. Neural interfaces add an inward-facing source to the observation stream by measuring neural activity directly. Such signals may eventually provide infor- mation about internal processes, including inner speech, perception, attention, or affect, that external observation cannot directly cap- ture. Existing demonstrations are much narrower: they decode attempted speech in a participant with paralysis [42], a small vocab- ulary of internal speech [40], and semantic content from perceived and imagined speech [36]. Neural signals could complement multi- modal capture by linking internal processes to the external situa- tions in which they occur. However, current systems provide neither unrestricted nor transparent access to thought, and open-ended access to thought remains the motivating limit case. Because such a channel could extend beyond deliberate disclosure, it heightens concerns about mental privacy, autonomy, identity, and control. 4.1 The Complementary Self-Model Self-tracking turns observations about daily life into persistent records that users can revisit. Personal informatics systems and the Quantified Self movement use such records to support self- knowledge [22]. Across 138 randomized studies, interventions de- signed to increase monitoring of goal progress improved goal at- tainment, with larger effects when information obtained through monitoring was recorded or reported [13]. In a personal health in- formatics study, participants used generative AI to analyze diverse tracking data, explore relationships between health measures and daily behavior, validate existing practices, and identify new health goals [4]. A personal AI can extend these practices by relating such signals across time and bringing them into planning and action, for example by weighing a poor night’s sleep against deadlines and project history before proposing a schedule. Over time, such a system can build a model that complements the user’s existing self-knowledge. Users know some aspects of their present experience firsthand, whereas an external system must infer those aspects from reports and observations. Repeated records can reveal changes over time that retrospective recall may miss [3]. For personal AI, these patterns may include recurring problems, shifts in time allocation, and gradual drift between schedules and stated goals. Some relevant knowledge is also difficult to put into words [30], limiting what manual reports can capture. Persistent memory can support this longer view by comparing reports and records across weeks or months. It can keep deadlines, recurring bottlenecks, and neglected projects in view, and can sur- face a sustained gap between stated priorities and recorded time allocation. Broader or more continuous channels can strengthen these comparisons: calendars and activity traces add evidence about scheduled and observed behavior, while multimodal capture pre- serves context absent from written reports. A sufficiently calibrated self-model could support counterfactual planning by estimating how candidate plans might interact with the user’s likely behavior, but blindly treating predicted behavior as an endorsed objective could reinforce habits the user wants to change, steer their choices, or justify actions without authorization. A complementary self-model can become a tool the user relies on for remembering, reflecting, and deciding, allowing it to func- tion as part of an extended cognitive process [6]. The user must retain control over what the system observes, remembers, and does. 4 Toward Personal Intelligence Through Cooperative ObservationPILA ’26 (KDD 2026), August 9, 2026, Jeju Island, Republic of Korea Phase 1 Manual reports + Stated goals, reflections – Effort, reporting gaps Phase 2 Behavioral context + Observed routines, actions – Intent remains hidden Phase 3 Multimodal capture + Situated multimodal context – Bystanders, data overload Phase 4 Neural interfaces + Unexpressed neural signals – Mental privacy, autonomy Broader, more continuous observation with less reliance on self-report Figure 2: Observation-channel spectrum for personal AI. Later phases can broaden the observation stream through more continuous modalities that rely less on self-report, while also increasing privacy, consent, and autonomy risks. The Good Regulator theorem provides a second motivation: ef- fective regulation depends on a model of what is being regulated [7]. For personal assistance, this points toward modeling the goals, commitments, and constraints relevant to the user’s decisions. 4.2 Task-Conditioned Selection and Compression Broader observation can leave the system with more history than it can use in any single decision. The big world hypothesis treats this mismatch as fundamental: an agent cannot fully perceive a world much larger than itself or represent the correct response to every state, so it must rely on approximation [15]. A personal AI must therefore select and compress its observation history for the task at hand. The information bottleneck formalizes the resulting trade-off: a compact representation should preserve information about the decision-relevant target while retaining as little of the full input as possible [38]. The relevant target changes across users, tasks, and time, and several targets may matter at once. Compression for one target can therefore remove evidence needed for another, so systems need task-specific summaries and, where the user’s policy allows it, a path back to the source material. Existing work asks whether retrieved context is sufficient to answer a query [16]. Personal AI poses a broader task-conditioned context allocation prob- lem: deciding how much user information to make available, from which sources, and in what representation and abstraction so that it supports the task within processing and disclosure constraints. Systems often seek broader coverage by combining channels with different formats and timescales. This can provide evidence for decisions that span several parts of the user’s life, but the sys- tem must decide which sources to consult, at what resolution, and how to reconcile observations collected on different timescales. An information-theoretic account of bounded rationality shows why hierarchy can help: trading expected utility against information- processing cost can induce abstractions at different levels of a de- cision hierarchy [10]. The Society of Mind provides a complemen- tary architectural view in which specialized processes interact to produce intelligence [23]. These perspectives suggest a recurring response to boundedness across the user-system boundary: The user externalizes experience into persistent records, and bounded agents select, compress, and route those records at task granularity. 5 Organizm: A Prototype for Cooperative Observation Organizm is a prototype for cooperative observation that has been used in practice over multiple months. It currently runs on manual reports, user-owned files, and lightweight calendar integration. Sec- tion 5.4 describes a six-month single-subject deployment, including changes in reporting frequency and planning inputs together with examples of user correction. Architecturally, Organizm implements the society of mind view [23] through a hierarchical multi-agent system coupled to a user-owned file hierarchy. Organizm uses hierarchical externalization and task-conditioned selection across the user-system boundary. The user externalizes personal context into persistent records organized at multiple gran- ularities. Because no single agent can use the full record at once, the system selects task-appropriate working sets and assigns spe- cialized agents to corresponding scopes. The coupled hierarchy therefore extends the user’s capacity while managing the system’s own capacity limits. It also combines file-native memory, coupled agent and information hierarchies, and a feedback loop that updates plans and governs progressive integration. 5.1 File-Native Memory and Index-First Traversal The system stores daily logs, project notes, and session summaries as ordinary files in a user-owned file hierarchy. This local-first design makes user control, data ownership, and software continuity architectural requirements while allowing the user to inspect, edit, or delete the memory directly [20]. Folder-level and project-level index files summarize what lives below them, allowing an agent to route through compressed rep- resentatives before reading full files [35]. An agent starts with a coarse index and follows it to more detailed files only when the task requires them. 5.2Coupled Agent and Information Hierarchies Within the system, different tasks require different context scopes. Immediate note processing does not need the full life history, while long-term reflection needs goals, values, and historical summaries. The system is based on a hierarchical multi-agent interface that is coupled to a hierarchical information store [35]. Specialized per- sonas operate at scopes ranging from immediate processing to long- term architecture. The agent hierarchy and the memory hierarchy are aligned by timescale and abstraction level (Figure 3). 5 PILA ’26 (KDD 2026), August 9, 2026, Jeju Island, Republic of KoreaYashar Talebirad, Osman Jime, Ali Parsaee, Eden Redman, Yongbin Kim, and Osmar R. Zaïane Long-term architecture Weekly strategy Daily coaching Immediate note processing Goals, values, slow summaries Project summaries, weekly reviews Daily logs, deadlines, inbox Session traces, new notes Agent personas Information tiers Priorities↓Summaries↑ Figure 3: The prototype couples agent personas to matching information tiers. Higher levels preserve abstraction, while lower levels keep current observations close to action. This coupling implements task-conditioned context allocation from Section 4.2: each agent initially loads the information tier matched to its task, while index-first traversal supplies finer detail when that tier is insufficient. Longer-horizon goals, values, and commitments guide project plans and daily tasks, while evidence from lower-level observations and outcomes flows upward through progressively coarser summaries to support reflection and revision. This aligns goal horizon with agent scope and memory granularity. The same design choice has been tested in an ML-engineering system, where tiered loading outperformed flat loading while using fewer output tokens [19]. 5.3 Feedback, Integration, and Control The prototype maintains daily plans, weekly reviews, and user corrections to improve future planning. Daily and weekly files compare planned activity against actual activity, turning self-report into a recurring observation channel. Cooperative observation here depends on what the user reports after acting. Useful plans give the user a reason to report outcomes, whereas generic plans can make reporting feel like overhead and weaken the loop. Corrections provide calibration: rejecting overam- bitious daily plans teaches the system how to plan, while feedback on anxiety-inducing reminders teaches deadline escalation. The same feedback loop can guide progressive integration across the observation-channel spectrum in Section 4. Organizm’s current integrations remain within Phases 1–2, as described in Section 5.4. Future additions (e.g., health summaries or wearable signals) should have a user-endorsed purpose and a visible control boundary. Later phases would also increase data volume and sensitivity, affecting where inference should run. High-volume streams may make remote inference expensive, while sensitive streams may favor local processing. A later version could assign each channel to a specialized agent that compresses its stream before broader planning and escalates when confidence is low, when a task crosses domains, or when the user requests broader review. The user would retain control over which channels are enabled. Table 1: Daily log files per month in the single-subject Orga- nizm deployment (2026). MonthJanFebMarAprMayJun Daily logs24912213130 5.4 Six-Month Deployment Sketch We report an illustrative sketch of one author’s first six months using Organizm (January–June 2026). 3 Manual logging varied early and became near daily by May–June (Table 1), while weekly re- views and monthly summaries were added over the same period. These changes are consistent with cooperative observation, but the retrospective single-subject record cannot identify their causes or support generalization; Section 6 describes the multi-subject studies needed to test the proposed feedback loop. Over the same period, the observation channel expanded within Phases 1–2 of the observation-channel spectrum without wear- ables or continuous capture. Calendar events entered the daily planning loop first, followed by a shared deadline registry, weekly and monthly review layers, and a finance channel loaded by a spe- cialized agent. Each addition was voluntary and tied to a specific purpose, with the resulting information kept in user-owned files that could be inspected or deleted. Corrections also changed the stored model and later plans: in one episode, the user corrected a cost model that omitted a recurring expense; in another, the system revised a daily plan after checking its interpretation of the user’s commitments against the live calendar. These corrections updated memory and future planning via the feedback loop in Section 5.3. 6 Evaluation Directions We propose three complementary studies to evaluate cooperative observation. An observation-channel ablation can test whether task-relevant context improves assistance with a fixed model. Ad- ditionally, a longitudinal deployment can test whether assistance quality affects what users later share, correct, or revoke. Finally, a self-model study can examine changes in the system’s knowledge and the user’s self-knowledge. Observation-channel ablations. To test whether observation im- proves assistance, an ablation study would hold the model, tools, prompt template, and sampling settings fixed. Before evaluation, researchers would assemble a suite of real, participant-specific tasks with checkable constraints or outcomes, grouped into predefined categories such as scheduling, prioritization, planning, and fore- casting. Each task would be evaluated under five separate context conditions: no personal context, a static profile, recent manual logs, behavioral or calendar traces, and a combined condition containing the profile, recent logs, and available traces. Prompt length would be matched where possible, with irrelevant context of similar length included as a control for the effect of additional tokens. Responses would be scored using criteria defined in advance for each task. When human judgment is needed, raters would see only the task and response. With graded channel configurations, the same design could test whether task performance changes smoothly, saturates, or exhibits task-specific thresholds as observation broadens. 3 The subject is an author and consented to publication of these aggregate records. 6 Toward Personal Intelligence Through Cooperative ObservationPILA ’26 (KDD 2026), August 9, 2026, Jeju Island, Republic of Korea For each scored claim in a response, the analysis would record whether its supporting information came from the task request, the context supplied in that condition, or an inference from other evidence. A claim would count as an inference only when the target fact is absent from the task and supplied context and can be checked against an independent record. The no-context condition measures performance from the task request alone. A provenance check records which supplied sources support the response and flags cases in which the target already appears in the input. Across conditions, the results would show which context sources improve which kinds of tasks. Usefulness and channel change. To test whether assistance affects future observation, a longitudinal deployment would record per- ceived usefulness and task performance, then track what the user shares, corrects, or revokes over subsequent interactions. For each source, the study would record whether the user adds it, keeps it unchanged, reduces the detail it provides, or removes it, since users may weigh benefit, reporting effort, and privacy cost differently across sources. Reporting these changes for each source preserves channel-specific patterns within the overall sharing trend. The study should also track changes in the types of tasks users bring to the system, since the feedback loop may change the task distribution the system observes. Complementary self-model evaluation. System knowledge can be evaluated through paired predictions of later, verifiable outcomes, with one prediction based on the task alone and another using the task plus prior personal records. User self-knowledge can be evaluated through changes in the accuracy of the user’s forecasts of time allocation and project delays, along with assessments of consistency between planned work and stated goals. 7 Safety and Governance Richer observation increases both what a personal AI can do and the consequences of misuse or failure. Safety depends on which inferences are retained, whether observation channels can be ma- nipulated, how feedback shapes behavior, how agents share au- thority, and whether users can withdraw access and retain their data. Because the person being modeled also supplies feedback and controls access, personal AI may provide a small-scale testbed for alignment under uncertainty about user values [12]. The framework distinguishes immediate approval, reflective benefit, and decisions about continued access. Observation privacy and integrity. Raw records can support sen- sitive conclusions. For instance, a user may knowingly share sleep data without anticipating that the system could infer a health con- dition. A breach can therefore expose both the source data and the conclusions drawn from it, and retained inferences should receive the same privacy protection as their sources. Continuous audio and visual systems also capture bystanders, whose consent cannot be supplied by the primary user. Neural and other high-bandwidth channels also separate per- mission to collect a signal from permission to derive a particular inference. Access to raw neural activity should not be treated as blanket authorization to decode inner speech, attention, affect, or other latent attributes. Controls should therefore be scoped to par- ticular purposes and types of inference, disclose what was inferred and retained, and support pausing, deletion, and revocation for both raw signals and derived representations. Furthermore, every added channel broadens the attack surface. Red-teaming work on autonomous agents shows that malicious content in the observation stream can redirect behavior even when the user’s instruction is sound [45]. 4 Shared skills create another path because their instructions enter the agent’s context and may request access to tools or data. 5 Rankings and download counts should not substitute for code review, signing, sandboxing, and revocable permissions. Finally, containment requires separate credentials and memory- write permissions, since a compromised component that can invoke every tool or rewrite shared memory can affect the entire system. Derived summaries should retain source provenance so that claims can be traced after compression. Execution transparency and auditability. User control also re- quires an inspectable record of system activity. A personal agent should record which components accessed personal information, invoked tools, delegated tasks, changed memory, or exercised per- missions, together with the authorization and outcome of each action. Execution provenance connects these records across evi- dence, tool use, memory, and agent actions, supporting auditing, failure diagnosis, and recovery [41]. Users should be able to review these traces, investigate failures, and confirm that revoked chan- nels or permissions are no longer used. Because traces can reveal sensitive information, their access, storage, and retention controls should reflect the sensitivity of the activity they record. Goals, feedback, and dependence. Goal attainment leaves open which goal should govern an action. Immediate requests can conflict with longer-term commitments, and actions that benefit the user can impose costs on others. A personal agent should be designed to surface these conflicts, seek the user’s direction on consequential trade-offs, and respect the consent of affected people. Once a governing goal is selected, the system must still infer from available feedback whether its actions advance that goal. Im- mediate approval can be an unreliable proxy, as optimizing it may favor flattery, avoidance of hard truths, or increased dependence. Deployed assistants already show sycophancy, and human pref- erence data can reward it [31]. Evaluation should therefore track delayed outcomes and dependence alongside immediate ratings, while inspection, correction, and revocation give users ways to respond when approval and benefit diverge. Preserving agency also requires attention to how delegation affects the user’s ability to revise goals, retain capabilities, and choose which parts of an activity to perform. 4 Reported failures include unauthorized bulk email deletion after an agent lost the user’s approval constraint during context compaction; see Kevin Okemwa, “Meta’s safety director handed OpenClaw AI agents the keys to her emails,” https://w.windowscentral.com/artificial-intelligence/meta-summer-yue- director-openclaw-ai-email-deletion (accessed July 2026). 5 Cisco researchers demonstrated a highly ranked malicious skill that combined direct prompt injection with silent data exfiltration; see Amy Chang and Vineeth Sai Narajala, “Personal AI Agents like OpenClaw Are a Security Nightmare,” https://blogs.cisco.com/ ai/personal-ai-agents-like-openclaw-are-a-security-nightmare (accessed July 2026). 7 PILA ’26 (KDD 2026), August 9, 2026, Jeju Island, Republic of KoreaYashar Talebirad, Osman Jime, Ali Parsaee, Eden Redman, Yongbin Kim, and Osmar R. Zaïane Multi-agent composition. Personal AI becomes a multi-agent sys- tem when agents delegate to one another, share memory or services, or interact with agents serving other users. Work on multi-agent safety distinguishes properties visible in isolation from interactive properties, including an agent’s capacity to influence others, sus- ceptibility to exploitation, modeling of other agents, and behavior around shared norms [37]. These properties depend on the coun- terparties and conditions in which agents interact, so evaluation must include varied agents and realistic threat models. Coordination among agents with complementary skills may also produce system-level capabilities and risks that single-agent evaluation misses. The distributional AGI safety framework ad- dresses this possibility through sandboxed agent economies with governed transactions, reputation management, and oversight [39]. Interacting personal models may also support inferences about communities, institutions, and people who never contributed data. Local control over each observation channel leaves these cross-user effects unresolved. Evaluation should therefore test components and their interactions under realistic configurations of delegation, permissions, shared memory, and external tool access. Revocation, portability, and storage. Control must remain avail- able after a channel is enabled. Users need to inspect stored memory, correct inferences, and revoke observation channels [21]; revoca- tion should propagate to dependent summaries, credentials, queued actions, backups, and synchronized copies. Users also need portable copies of data gathered by their devices so they can change soft- ware as needed without losing their history. Restrictions in health and wearable ecosystems limit this control, and discontinued hard- ware can strand data in vendor-controlled channels. 6 Open APIs, documented formats, and user-controlled stores reduce this risk, while local-first and privacy-aware architectures keep ordinary use less dependent on a provider [20,21]. Keeping memory local concentrates security risk on the user’s devices, although remote storage can reduce dependence on any one device and support durable backups and synchronization, with every copy subject to equivalent access and deletion protections. Remote retrieval introduces another privacy problem because ac- cess patterns can reveal private information even when records are protected. Oblivious RAM (ORAM) hides memory-access patterns [24]; Opal applies this approach to personal AI memory so that the storage provider cannot learn which records a query retrieves [18]. These systems illustrate how storage architecture shapes privacy when personal memory extends beyond the user’s device. 8 Limitations Several limitations follow from the framework’s current stage. Cooperative observation depends on consistent user feedback, which can be burdensome to provide. Privacy concerns can also limit what users share, and outcomes may depend on several active observation channels. A failure may therefore provide little guid- ance about which channel needs correction. Sparse or inconsistent feedback also makes channel changes harder to interpret. Future 6 Recent examples include the Limitless Pendant, which ceased sale when Meta ac- quired the company in Dec. 2025 (https://techcrunch.com/2025/12/05/meta-acquires- ai-device-startup-limitless/, accessed July 2026), and Humane’s AI Pin, whose servers shut down days after HP’s acquisition in Feb. 2025 (https://techcrunch.com/2025/02/18/ humanes-ai-pin-is-dead-as-hp-buys-startups-assets-for-116m/, accessed July 2026). studies can combine channel-level provenance with controlled ab- lations to isolate these effects. Reflective benefit푈 푡 , trust, and task relevance require operational definitions specific to the user, task, and evaluation horizon. The framework also identifies task-conditioned context allocation as a design problem. Comparing allocation methods under processing and disclosure constraints is a separate empirical agenda. The delegated task distribution may change as users learn which tasks the system handles well. This feedback between system be- havior and later inputs is related to performative prediction [28] and complicates performance comparisons over time. The behavioral effect of Organizm’s coupled agent and memory hierarchies remains unevaluated relative to flat memory or a single agent operating over the same files. Future ablations should test whether the coupling helps and by how much. Finally, the six-month deployment sketch covers one user under naturalistic conditions. The observed reporting frequency, channel additions, and corrections do not identify causal effects or support generalization, which require controlled multi-user studies. 9 Conclusion Personal agents depend on context relevant to the user’s goals, yet much of that context remains outside the system even as de- vices capture richer traces of daily life. As observation grows richer, task-conditioned context allocation determines which sources and representations enter a decision under processing and disclosure constraints. Cooperative observation describes the resulting feed- back loop: the system acts on what it can observe, the user evaluates the result, and that evaluation shapes what the system may observe next. Whether the user maintains a channel depends on the system’s usefulness and trustworthiness, the channel’s effort and privacy costs, and the control the user retains. Evaluation should examine whether task-relevant context im- proves assistance, how assistance quality affects later sharing, cor- rection, and revocation, and how the system’s knowledge and the user’s self-knowledge change over time. Organizm implements the loop through manual reports, user- owned memory, and iterative planning. Its six-month deployment records changes in reporting and planning inputs. The single-user record cannot determine why these changes occurred; controlled multi-user studies are needed to test these causes. As sensing be- comes more continuous, open software, local-first memory, portable data, and auditable execution can support user control. Future agents should apply personal knowledge to user-chosen goals within user-set delegation boundaries. Users should see and control what personal AI observes, infers, remembers, delegates, and does. Acknowledgments This research was supported by the Alberta Machine Intelligence Institute (Amii) and the Canada CIFAR AI Chairs Program. We also thank colleagues at the Network for Applied Technology (NAT), including Yash Mouje, Liliya Eghdamian, Qendrim Beka, and Eric Fung, for their comments and support. References [1] Song Bian, Tao Yu, Shivaram Venkataraman, and Youngsuk Park. 2026. Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs. In International 8 Toward Personal Intelligence Through Cooperative ObservationPILA ’26 (KDD 2026), August 9, 2026, Jeju Island, Republic of Korea Conference on Learning Representations. https://arxiv.org/abs/2510.18245 [2]Sophie C. Boerman, Sanne Kruikemeier, and Frederik J. Zuiderveen Borgesius. 2017. Online Behavioral Advertising: A Literature Review and Research Agenda. Journal of Advertising 46, 3 (2017), 363–376. doi:10.1080/00913367.2017.1339368 [3]Niall Bolger, Angelina Davis, and Eshkol Rafaeli. 2003. Diary Methods: Capturing Life as It Is Lived. Annual Review of Psychology 54 (2003), 579–616. doi:10.1146/ annurev.psych.54.101601.145030 [4]Shaan Chopra, Katherine Juarez, James Fogarty, and Sean A. Munson. 2025. En- gagements with Generative AI and Personal Health Informatics: Opportunities for Planning, Tracking, Reflecting, and Acting around Personal Health Data. Pro- ceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 9, 3 (2025), 1–33. doi:10.1145/3749503 [5] Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep Reinforcement Learning from Human Preferences. In Ad- vances in Neural Information Processing Systems 30. https://proceedings.neurips. c/paper/2017/hash/d5e2c0adad503c91f91df240d0cd4e49-Abstract.html [6]Andy Clark and David J. Chalmers. 1998. The Extended Mind. Analysis 58, 1 (1998), 7–19. doi:10.1093/analys/58.1.7 [7]Roger C. Conant and W. Ross Ashby. 1970. Every Good Regulator of a System Must Be a Model of That System. International Journal of Systems Science 1, 2 (1970), 89–97. doi:10.1080/00207727008920220 [8] Thomas M. Cover and Joy A. Thomas. 2006. Elements of Information Theory (2nd ed.). Wiley-Interscience, Hoboken, NJ. doi:10.1002/047174882X [9]Jim Gemmell, Gordon Bell, and Roger Lueder. 2006. MyLifeBits. Commun. ACM 49, 1 (2006), 88–95. doi:10.1145/1107458.1107460 [10] Tim Genewein, Felix Leibfried, Jordi Grau-Moya, and Daniel Alexander Braun. 2015. Bounded Rationality, Abstraction, and Hierarchical Decision-Making: An Information-Theoretic Optimality Principle. Frontiers in Robotics and AI 2 (2015). doi:10.3389/frobt.2015.00027 [11] C. A. E. Goodhart. 1984. Problems of Monetary Management: The UK Experience. In Monetary Theory and Practice. Macmillan Education UK, London, 91–121. doi:10.1007/978-1-349-17295-5_4 [12]Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell. 2016. Cooperative Inverse Reinforcement Learning. In Advances in Neural Informa- tion Processing Systems 29. https://proceedings.neurips.c/paper/2016/hash/ c3395d46c34fa7fd8d729d8cf88b7a8-Abstract.html [13] Benjamin Harkin, Thomas L. Webb, Betty P. I. Chang, Andrew Prestwich, Mark Conner, Ian Kellar, Yael Benn, and Paschal Sheeran. 2016. Does Monitoring Goal Progress Promote Goal Attainment? A Meta-Analysis of the Experimental Evidence. Psychological Bulletin 142, 2 (2016), 198–229. doi:10.1037/bul0000025 [14] Steve Hodges, Lyndsay Williams, Emma Berry, Shahram Izadi, James Srinivasan, Alex Butler, Gavin Smyth, Narinder Kapur, and Ken Wood. 2006. SenseCam: A Retrospective Memory Aid. In UbiComp 2006: Ubiquitous Computing (Lecture Notes in Computer Science, Vol. 4206). Springer, 177–193. doi:10.1007/11853565_11 [15]Khurram Javed and Richard S. Sutton. 2024. The Big World Hypothesis and its Ramifications for Artificial Intelligence. In Finding the Frame: An RLC Workshop. https://openreview.net/forum?id=Sv7DazuCn8 [16]Hailey Joren, Jianyi Zhang, Chun-Sung Ferng, Da-Cheng Juan, Ankur Taly, and Cyrus Rashtchian. 2025. Sufficient Context: A New Lens on Retrieval-Augmented Generation Systems. In International Conference on Learning Representations. https://openreview.net/forum?id=Jjr2Odj8DJ [17]Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra. 1998. Planning and Acting in Partially Observable Stochastic Domains. Artificial Intelligence 101, 1–2 (1998), 99–134. doi:10.1016/S0004-3702(98)00023-X [18]Darya Kaviani, Alp Eren Ozdarendeli, Jinhao Zhu, Yu Ding, and Raluca Ada Popa. 2026. Opal: Private Memory for Personal AI. https://arxiv.org/abs/2604.02522 [19]Yongbin Kim, Yashar Talebirad, and Osmar R. Zaiane. 2026. Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering. https: //arxiv.org/abs/2606.30911v2 Version 2. [20]Martin Kleppmann, Adam Wiggins, Peter van Hardenberg, and Mark Mc- Granaghan. 2019. Local-First Software: You Own Your Data, in Spite of the Cloud. In Proceedings of the 2019 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software. 154–178. doi:10.1145/3359591.3359737 [21]Marc Langheinrich. 2001. Privacy by Design: Principles of Privacy-Aware Ubiqui- tous Systems. In UbiComp 2001: Ubiquitous Computing (Lecture Notes in Computer Science, Vol. 2201). Springer, 273–291. doi:10.1007/3-540-45427-6_23 [22]Ian Li, Anind K. Dey, and Jodi Forlizzi. 2010. A Stage-Based Model of Personal Informatics Systems. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 557–566. doi:10.1145/1753326.1753409 [23] Marvin Minsky. 1986. The Society of Mind. Simon & Schuster, New York. [24]Rafail Ostrovsky. 1990. Efficient Computation on Oblivious RAMs. In Proceedings of the Twenty-Second Annual ACM Symposium on Theory of Computing. ACM, 514–523. doi:10.1145/100216.100289 [25]Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training Language Models to Follow Instructions with Human Feedback. In Advances in Neural Information Processing Systems 35. 27730–27744. doi:10.52202/068431-2011 [26]German I. Parisi, Ronald Kemker, Jose L. Part, Christopher Kanan, and Stefan Wermter. 2019. Continual Lifelong Learning with Neural Networks: A Review. Neural Networks 113 (2019), 54–71. doi:10.1016/j.neunet.2019.01.012 [27] Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–22. doi:10.1145/3586183.3606763 [28] Juan Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, and Moritz Hardt. 2020. Performative Prediction. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119). PMLR, 7599–7609. https://proceedings.mlr.press/v119/perdomo20a.html [29]PIN AI Team. 2025. GOD model: Privacy Preserved AI School for Personal Assistant. https://arxiv.org/abs/2502.18527 [30]Michael Polanyi. 1966. The Tacit Dimension. Doubleday, Garden City, NY. Terry Lectures, Yale University. [31]Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2024. Towards Understanding Sycophancy in Language Models. In The Twelfth International Conference on Learning Representations. https://arxiv.org/abs/2310. 13548 [32] Jonathan Stray. 2020. Aligning AI Optimization to Community Well-Being. International Journal of Community Well-Being 3, 4 (2020), 443–463. doi:10.1007/ s42413-020-00086-3 [33] Jonathan Stray, Alon Halevy, Parisa Assar, Dylan Hadfield-Menell, Craig Boutilier, Amar Ashar, Chloe Bakalar, Lex Beattie, Michael Ekstrand, Claire Leibowicz, Connie Moon Sehat, Sara Johansen, Lianne Kerlin, David Vickrey, Spandana Singh, Sanne Vrijenhoek, Amy Zhang, McKane Andrus, Natali Helberger, Polina Proutskova, Tanushree Mitra, and Nina Vasan. 2024. Building Human Values into Recommender Systems: An Interdisciplinary Synthesis. ACM Transactions on Recommender Systems 2, 3 (2024), 1–57. doi:10.1145/3632297 [34] Jonathan Stray, Ivan Vendrov, Jeremy Nixon, Steven Adler, and Dylan Hadfield- Menell. 2021. What are you optimizing for? Aligning Recommender Systems with Human Values. https://arxiv.org/abs/2107.10939 [35]Yashar Talebirad, Ali Parsaee, Csongor Y. Szepesvári, Amirhossein Nadiri, and Osmar R. Zaïane. 2026. Toward a Theory of Hierarchical Memory for Language Agents. https://arxiv.org/abs/2603.21564 ICLR 2026 Workshop on Memory for LLM-Based Agentic Systems. [36] Jerry Tang, Amanda LeBel, Shailee Jain, and Alexander G. Huth. 2023. Semantic Reconstruction of Continuous Language from Non-Invasive Brain Recordings. Nature Neuroscience 26 (2023), 858–866. doi:10.1038/s41593-023-01304-9 [37]Cecilia Elena Tilli. 2026. Agent Properties for Multi-Agent Safety. In ICLR 2026 Workshop on Agents in the Wild. https://openreview.net/forum?id=00KtDBoKnO [38]Naftali Tishby, Fernando C. Pereira, and William Bialek. 1999. The Information Bottleneck Method. In Proceedings of the 37th Annual Allerton Conference on Communication, Control, and Computing. 368–377. https://arxiv.org/abs/physics/ 0004057 [39] Nenad Tomašev, Matija Franklin, Julian Jacobs, Sébastien Krier, and Simon Osin- dero. 2025. Distributional AGI Safety. https://arxiv.org/abs/2512.16856v2 Version 2, revised May 19, 2026. [40]Sarah K. Wandelt, David A. Bjånes, Kelsie Pejsa, Brian Lee, Charles Liu, and Richard A. Andersen. 2024. Representation of Internal Speech by Single Neurons in Human Supramarginal Gyrus. Nature Human Behaviour 8 (2024), 1136–1149. doi:10.1038/s41562-024-01867-y [41]Yiqi Wang, Jiaqi Zhang, Taotao Cai, Zirui Liu, Qingqiang Sun, Zequn Sun, Zhangkai Wu, Manqing Dong, Mingkai Zheng, Xuefei Yin, and Yanming Zhu. 2026. From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents. https://arxiv.org/abs/2606.04990 [42]Francis R. Willett, Erin M. Kunz, Chaofei Fan, Donald T. Avansino, Guy H. Wilson, Eun Young Choi, Foram Kamdar, Matthew F. Glasser, Leigh R. Hochberg, Shaul Druckmann, Krishna V. Shenoy, and Jaimie M. Henderson. 2023. A High- Performance Speech Neuroprosthesis. Nature 620, 7976 (2023), 1031–1036. doi:10. 1038/s41586-023-06377-x [43]Chaojun Xiao, Jie Cai, Weilin Zhao, Biyuan Lin, Guoyang Zeng, Jie Zhou, Zhi Zheng, Xu Han, Zhiyuan Liu, and Maosong Sun. 2025. Densing Law of LLMs. Nature Machine Intelligence 7 (2025), 1823–1833. doi:10.1038/s42256-025-01137-0 [44]Yan Xu, Brennan Jones, Hannah Nguyen, Qisheng Li, and Stefan Scherer. 2025. From Goals to Actions: Designing Context-Aware LLM Chatbots for New Year’s Resolutions. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). Association for Computing Machinery, New York, NY, USA, Article 56, 17 pages. doi:10.1145/3719160.3736637 [45]Haochen Zhao and Shaoyang Cui. 2026. ClawTrap: A MITM-Based Red-Teaming Framework for Real-World OpenClaw Security Evaluation. https://arxiv.org/ abs/2603.18762 9