Paper deep dive
Physiological World Models for Human State Transitions
Chongyang Zhang, Rendong Wang, Hao Zheng, Hanwen Zhang, Yang Liu, Xiaolong Wei, Bin Chong
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/18/2026, 6:10:56 AM
Summary
The paper proposes the Physiological World Model (PWM), an event-conditioned framework for modeling how an individual's integrated physiological state transitions in response to real-world events, behaviors, contexts, and interventions. It introduces the HumanState Transition Token as a structured unit for training and evaluation, defines a four-level capability ladder (L1-L4) ranging from state representation to bounded intervention planning, and outlines four data acquisition protocols (P1-P4) and six benchmark tasks to advance personalized health management and clinical decision support.
Entities (13)
Relation Signals (10)
HumanState Transition Token → contains → HumanState
confidence 95% · HumanStateTransitionToken = HumanState t , WorldEvent t , HumanAction t , Context t , Intervention t , HumanState t+∆t , Outcome t+∆t , DataQuality t:t+∆t
HumanState Transition Token → contains → WorldEvent
confidence 95% · HumanStateTransitionToken = HumanState t , WorldEvent t , HumanAction t , Context t , Intervention t , HumanState t+∆t , Outcome t+∆t , DataQuality t:t+∆t
HumanState Transition Token → contains → HumanAction
confidence 95% · HumanStateTransitionToken = HumanState t , WorldEvent t , HumanAction t , Context t , Intervention t , HumanState t+∆t , Outcome t+∆t , DataQuality t:t+∆t
HumanState Transition Token → contains → Intervention
confidence 95% · HumanStateTransitionToken = HumanState t , WorldEvent t , HumanAction t , Context t , Intervention t , HumanState t+∆t , Outcome t+∆t , DataQuality t:t+∆t
Physiological World Model → uses → HumanState Transition Token
confidence 95% · We introduce the HumanState Transition Token, a structured, quality-scored unit that connects the physiological state before an event with the event or action...
Physiological World Model → supports → Capability Level L1
confidence 90% · Large-scale unlabelled or weakly labelled recordings can support L1 state-representation learning.
Physiological World Model → supports → Capability Level L4
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical visits. However, most health artificial intelligence systems are designed to recognize current states, estimate risks or analyse individual biomarkers. They do not directly model how physiological states change in response to real-world events, behaviours, contexts and interventions. Here we propose the Physiological World Model (PWM), an event-conditioned framework for learning these changes at the level of the whole person. We introduce the HumanState Transition Token, a structured, quality-scored unit that connects the physiological state before an event with the event or action, relevant context and intervention information, the physiological trajectory after the event, observed outcomes and data quality. We describe four capability levels, from state representation to bounded intervention planning, together with four data acquisition and validation protocols. We also propose six benchmark tasks covering HumanState representation, forecasting across multiple timescales, individualized response prediction, simulation of alternative interventions, bounded planning and reliability under distribution shift. Together, this framework provides a practical path towards personalized health management, behavioural intervention design and clinician-supervised decision support, while clearly separating prediction from causal inference and making uncertainty, safety, governance and limits of use explicit.
Tags
Links
- Source: https://arxiv.org/abs/2608.15309v1
- Canonical: https://arxiv.org/abs/2608.15309v1
Trouble viewing inline? Open PDF directly →
Full Text
92,437 characters extracted from source content.
Expand or collapse full text
Perspective Physiological world models for human state transitions Chongyang Zhang 1† , Rendong Wang 1† , Hao Zheng 1 , Hanwen Zhang 1 , Yang Liu 2 , Xiaolong Wei 1 , Bin Chong 2* 1 Fullive-AI. 2 Peking University. † Equal contribution, ∗ Corresponding author Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical visits. However, most health artificial intelligence systems are designed to recognize current states, estimate risks or analyse individual biomarkers. They do not directly model how physiological states change in response to real-world events, behaviours, contexts and interventions. Here we propose the Physiological World Model (PWM), an event-conditioned framework for learning these changes at the level of the whole person. We introduce the HumanState Transition Token, a structured, quality-scored unit that connects the physiological state before an event with the event or action, relevant context and intervention information, the physiological trajectory after the event, observed outcomes and data quality. We describe four capability levels, from state representation to bounded intervention planning, together with four data acquisition and validation protocols. We also propose six benchmark tasks covering HumanState representation, forecasting across multiple timescales, individualized response prediction, simulation of alternative interventions, bounded planning and reliability under distribution shift. Together, this framework provides a practical path towards personalized health management, behavioural intervention design and clinician-supervised decision support, while clearly separating prediction from causal inference and making uncertainty, safety, governance and limits of use explicit. Keywords: Physiological World Models, Event-conditioned Modeling, Causal Inference, Counterfactual Simulation, Synchronized Multimodal Data Protocol, Digital Health 1 Introduction Observing human physiology continuously in the real world, rather than episodically in the clinic, represents one of the most consequential shifts in personal health monitoring since the invention of the Holter monitor [1]. Over the past decade, the integration of miniaturized sensors, ubiquitous wireless connectivity, and cloud-scale data infrastructure has brought multimodal physiological monitoring to daily life on an unprecedented scale [1,2]. Smartwatches and fitness bands capture cardiovascular, activity, temperature, and electrodermal signals while estimating sleep patterns [2–5]; continuous glucose monitors and wearable ECGs extend metabolic and cardiac monitoring beyond the clinic [6–8]; and smartphones with environmental sensors record contextual factors such as light, noise, temperature, and location, which are increasingly integrated with physiological time series [9–12]. However, the infrastructure for modelling and reasoning over these data has not kept pace with data collection [13,14]. Physiological records are generated and organized around devices, care settings, disease categories, and institutional silos, and are typically treated as modality- or task-specific features rather than as coordinated representations of an evolving human state [15,16]. As a result, real-world events, individual behaviours, physiological responses, and intervention outcomes are rarely aligned along a common temporal axis. A unified modelling framework is therefore needed to integrate multimodal observations and characterize how an individual’s physiological state evolves in response to specific events and actions in defined contexts [10, 17–19]. Artificial intelligence has been increasingly applied to health data. Yet most current systems are designed for state recognition, anomaly detection, or risk prediction. They answer what an individual’s current physiological state is, whether a biomarker is abnormal, or whether a future clinical outcome is likely [20–24]. Although these functions have demonstrated clinical utility [8,23], current systems often model physiology through isolated observations or input–output associations, rather than as a dynamic process shaped by events, behaviours, and context [14,25]. What remains insufficiently addressed is how physiological states change over time: why a state changes, why the same event produces different responses across individuals, and how a trajectory might differ under alternative actions. Addressing this gap requires models that learn event-conditioned physiological state transitions, rather than merely recognizing current states or predicting predefined outcomes [13, 14, 25, 26]. arXiv:2608.15309v1 [cs.AI] 15 Aug 2026 Figure 1 Physiological world models within a multiscale hierarchy of biological and health modelling. The proposed taxonomy orders modelling scopes from population and One Health systems [37], through integrated physiology at the individual level, to tissue and organ [38], cellular [39], and molecular or biomolecular processes [40]. PWMs occupy the individual-level layer, whereas models at broader scales address population, cross-species, and environmental dynamics, and models at finer scales resolve processes within organs, cells, or molecules. The lower panel shows illustrative domain-specific specializations within the PWM level, including EEG, ECG/PPG [41], sleep, metabolic, physiological-indicator, and disease-trajectory WMs [35, 36, 42]. These subdomains include both existing model instances and prospective directions. Arrows indicate the ordering of modelling scopes rather than causal or computational flow. Recent advances in world modelling provide a conceptual basis for extending world-model principles to human physiology. World models learn task-relevant state representations and transition dynamics that support prediction, simulation, and planning [25,27,28]. They have advanced game-playing agents, robotic control, and autonomous navigation [25,29–31], and related approaches are beginning to appear in biomedical research, including molecular and cellular modelling, physiological-signal modelling, and clinical trajectory analysis [26, 32–36]. However, existing biomedical world-model research has rarely centred on how an individual’s macro-level physiology evolves in response to real-world events and interventions. Modelling physiology at this scale requires temporally aligning multimodal physiological observations with real-world events, actions, interventions, and contexts across multiple timescales [5,11, 12]. In this Perspective, we define the Physiological World Model (PWM) as an event-conditioned framework for modelling transitions in an individual’s macro-level physiological state. Whereas digital twins generally emphasize individual- specific virtual representations and continuous state synchronization [43,44], PWMs focus on learning event-conditioned physiological state-transition dynamics; we discuss this distinction further in Section 6.6. As illustrated in Figure 1, our proposed taxonomy positions PWMs at the level of integrated individual physiology, between models of broader popula- tion and One Health systems and finer-scale models of tissues, organs, cells, and molecules [37–39,45]. To support the modelling of physiological state transitions, we introduce the HumanState Transition Token as a structured, quality- scored unit for transition learning and evaluation. We then describe an architecture-agnostic learning framework and an L1–L4 PWM capability ladder extending from state representation to bounded intervention planning. Four levels of data acquisition and validation (P1–P4) provide progressively stronger forms of evidence for these capabilities, although the correspondence is not one-to-one. We further propose six complementary benchmark tasks covering HumanState repre- sentation, multi-horizon forecasting, event-conditioned and individualized response prediction, alternative intervention simulation, bounded planning, and reliability under distribution shift. These tasks are linked to near-term personal health applications, mid-term behavioural intervention design, and longer-term clinical decision support, with progressively stronger validation and governance requirements [23–25, 46–48]. 2 Defining Physiological World Models 2.1 Background and Scope World models have recently emerged as a class of approaches for learning predictive representations of dynamical systems in artificial intelligence and reinforcement learning [27–30]. Rather than reconstructing an environment in 2 full, they learn task-relevant state representations and transition dynamics that support forecasting, simulation, and planning [27,29,30,49]. World-model approaches have begun to appear in clinical trajectory modelling and decision support [26], while related predictive-representation methods are being explored in wearable biosignal modelling [33]. However, approaches that connect real-world events, human actions, contexts, and interventions to transitions in an individual’s integrated physiological state remain underexplored. From this perspective, we define the Physiological World Model (PWM) as a human-centred, event-conditioned model of the transition dynamics of an individual’s integrated physiological state. Here, human-centred means that the evolving physiological state of the whole person, situated within their physical, social and behavioural environment, is the primary object of modelling. A PWM therefore focuses on macro-scale, temporally evolving physiology rather than attempting to reconstruct human biology in its entirety or being limited to episodic clinical records. Its distinguishing objective is to connect internal physiological dynamics with external events, actions, contexts, and interventions. This enables future state trajectories to be predicted and, when supported by appropriate interventional evidence, compared across alternative actions or interventions. A PWM does not treat physiological and contextual data as a flat collection of features. Instead, it represents phys- iological information across three levels while retaining contextual measurements as distinct conditioning variables. The Raw Observation Layer contains physiological sensor streams and time-aligned contextual measurements. The Interpretable Physiology Layer transforms physiological observations into interpretable intermediate variables, such as autonomic regulation, metabolic stability, circadian phase, sleep pressure, and recovery capacity [6,50–53]. The Latent HumanState Layer integrates these physiological observations and intermediate variables into an individual’s macro-level physiological state. Together, these layers connect heterogeneous measurements to a unified state representation without assuming that all devices measure the same variables or provide the same level of fidelity. The resulting latent state supports transition prediction, simulation, and intervention comparison. Having defined what a PWM represents, we next formalize how that state changes over time. 2.2 HumanState and Transition Dynamics Beyond representing an individual’s current physiological state, a PWM models how that state evolves in response to events, actions and interventions within defined contexts. Standard stochastic world models commonly represent action- conditioned dynamics through a transition distribution of the formp(s t+1 | s t , a t )[25]. We extend this formulation by incorporating observed context and an explicit prediction horizon: p s t+∆t | s t , a t , c t , ∆t .(1) Here,s t ands t+∆t denote the current and future states, respectively;a t denotes an action;c t denotes contextual information;tdenotes the current time; and∆tdenotes the prediction horizon. Extending this formulation to human physiology, a PWM models the conditional distribution of future HumanState as: p HumanState t+∆t | HumanState t , WorldEvent t , HumanAction t , Context t , Intervention t , ∆t . (2) This formulation aligns the PWM with a general world-model transition formulation while making the drivers of physiological change explicit. It allows the model to represent multiple plausible future trajectories and their associated uncertainty via multi-step rollouts, rather than returning only a deterministic point estimate. Comparisons across alternative interventions should be interpreted as causal effects only when supported by appropriate study designs and assumptions. This formulation specifies the conceptual dependencies governing physiological state transitions, without prescribing how the variables are encoded or which neural architecture is used to model them. The terms in the PWM formulation are defined as follows. HumanState t denotes the latent macro-level physiological state of an individual at timet. It is inferred from an integra- tion of physiological observations such as heart rate, blood glucose, blood pressure, oxygen saturation, body temperature, and validated sleep-related measures with relevant disease-related status, including chronic conditions, comorbidities, and disease stage, where available. Rather than corresponding to any single observed measure,HumanState t reflects the individual’s overall physiological condition and the accumulated effects of prior physiological changes, behaviours, and treatments up to timet. Where reliable assessments are available, subjective variables such as mood and perceived stress may also inform the inference of this state [54]. 3 WorldEvent t captures external events and exposures that may influence HumanState. Examples include abrupt changes in environmental noise, ambient temperature, or light exposure; traffic disruptions; salient social interactions; and unexpected work-related demands.HumanAction t represents observed volitional behaviours, such as eating, exercise, caffeine or alcohol intake, medication adherence, screen use, and sleep timing. Context t specifies the observed individual and situational information available at or before timetthat conditions a state transition but is not itself represented as the contemporaneous HumanState, focal WorldEvent, HumanAction, or Intervention. It may include time-invariant or slowly varying individual background, such as age, body size, past medical history, preferences, and habitual behaviours, as well as time-varying pre-transition context, such as time of day, day of week, previous-night sleep quality, and recent behavioural or dietary history. Actual eating, exercise, and other focal behaviours should be represented as HumanAction, whereas current physiological measures and current disease status belong to HumanState. Intervention t refers to a deliberately assigned or planned modification intended to alter a physiological trajectory. Examples include a dietary regimen, exercise prescription, sleep-schedule adjustment, medication change, or behavioural modification programme. This definition distinguishes a planned intervention from a naturally occurring or merely observed action, even when the two involve the same behaviour. Finally,HumanState t+∆t represents the future physiological state at the specified prediction horizon∆t. A physi- ological trajectory can be obtained by rolling out the transition model across multiple time steps. The horizon may range from minutes to years, depending on the task. Postprandial glucose responses and heart-rate recovery require short-horizon prediction, whereas physiological adaptations to sustained training, chronic disease progression, and other long-term physiological changes require substantially longer horizons [55,56]. Different horizons may also require different temporal resolutions or hierarchical transition components. 2.3 HumanState Transition Token Continuous recordings provide the observational foundation for PWM development, but transition learning requires event-anchored samples constructed from these streams. The state transition is therefore the fundamental modelling object of a PWM, while a HumanState Transition Token is the corresponding structured data instance for training and evaluation: HumanStateTransitionToken = HumanState t , WorldEvent t , HumanAction t , Context t , Intervention t , HumanState t+∆t , Outcome t+∆t , DataQuality t:t+∆t . (3) A transition token is not an arbitrary time window extracted from a continuous recording. It is anchored to an identifiable event, behaviour, or intervention around which a physiological response can be evaluated. Each token includes a pre- event baseline, a post-event response window, contextual variables, intervention information when applicable, and an observed outcome (for temporal descriptions such as pre-event and post-event windows, “event” denotes the focal transition anchor, which may be a WorldEvent, HumanAction, or Intervention). Fields that are not applicable to a given transition should be explicitly encoded as absent. In Eq.(3),HumanState t andHumanState t+∆t denote inferred macro-level physiological states derived from the corresponding pre- and post-event observation windows, rather than directly observed measurements. By contrast, Outcome t+∆t denotes observable physiological readouts and task-specific endpoints at the evaluation horizon. Where an observation decoder is used, it maps the predicted state representation to \ Outcome t+∆t , which can be compared with the observedOutcome t+∆t to assess the predicted transition. For example, pre-exercise heart rate, heart-rate variability, and glucose measurements may informHumanState t ; post-exercise observations may informHumanState t+∆t , while the observed heart-rate recovery curve serves as an Outcome.DataQuality t:t+∆t summarizes the completeness, temporal alignment, sensor quality, annotation reliability, and overall confidence of the token across the full interval. These elements allow the reliability of each transition sample to be assessed explicitly. This definition changes the scaling logic of PWM development. Traditional deep-learning approaches to physiological signals often treat scaling as the accumulation of larger volumes of continuous recordings [57,58]. For transition learning, however, the most informative samples are high-confidence HumanState Transition Tokens. A dataset may contain millions of hours of single-modality wearable data yet yield few usable transition tokens when event annotations 4 Figure 2 An illustrative functional architecture for a Physiological World Model. Event, action, HumanState, context, and intervention are encoded through specific encoders and taken as conditional input to a latent-state transition model; rollout branches represent alternative future trajectories and uncertainty, and task-specific heads map predicted states to observable physiological outcomes. and contextual labels are absent. By contrast, a smaller dataset may provide richer supervision when events, contexts, pre- and post-event physiological windows, interventions, and outcomes are carefully documented. The role of transition tokens changes across capability levels. Large-scale unlabelled or weakly labelled recordings can support L1 state-representation learning. Progress from L2 state-transition prediction to L3 counterfactual simulation and L4 intervention planning increasingly depends on transition tokens with reliable event, action, context, intervention, outcome, and quality information. PWM development therefore requires not only more data, but also more diverse, higher-quality, and more information-dense transition evidence. This scaling logic motivates the P1–P4 framework for data acquisition and validation. P1 characterizes baseline physiological distributions, P2 introduces structured event and context annotation, P3 adds assigned interventions and repeated observations, and P4 provides prospective shadow-mode validation and adaptive empirical feedback. In summary, the basic unit for learning PWM transition dynamics is not a biomarker value at a single time point. It is a structured and verifiable record of how an individual’s physiological state changes under specific events, actions, and interventions within defined contexts. 3 Model Architecture and Learning Paradigm Section 2 defines what is represented by a PWM and the structure of HumanState Transition Tokens. Translating this definition into a trainable system requires both a latent state representation and a conditional transition model. Figure 2 presents one possible functional architecture for such a system, informed by latent-dynamics and JEPA-style predictive learning [25,59,60]. In this illustrative architecture, the core components include a state encoder, a conditional transition model, an uncertainty-estimation component, and mechanisms for encoding events, actions, interventions, and context. Observation decoders or task-specific outcome heads may be added when required to map predicted state representations to observable outcomes. Counterfactual rollout and planning components are introduced at higher capability levels rather than treated as minimum requirements. These components can be implemented using different neural architectures. The schematic specifies functional requirements rather than prescribing a particular neural architecture. The conditional transition distribution defined in Section 2 can be parameterized in latent space as: p θ s t+∆t | s t , e t , a t , i t , c t , ∆t ,(4) where s t and s t+∆t are model-specific latent representations of HumanState t and HumanState t+∆t , respectively, e t is an encoded representation ofWorldEvent t ,a t is an encoded representation ofHumanAction t ,i t is an encoded representation of a deliberately assigned or plannedIntervention t , andc t is an encoded representation ofContext t . The choice of horizon∆tdepends on the physiological process, and different horizons may require different temporal resolutions or hierarchical transition components [26, 61–64]. Large-scale physiological observations can support state-representation learning, whereas HumanState Transition Tokens provide event-anchored supervision for learning conditional transition dynamics. Learning these representations and 5 dynamics from real-world data presents several challenges: physiological observations are heterogeneous and irregularly sampled, events and interventions are sparsely and inconsistently labelled, and individual baselines vary substantially. The learning paradigms reviewed below offer complementary strategies for addressing these challenges. 3.1 Learning Paradigms for PWMs Given the functional definition above, the remaining question is how to learn a robust and temporally coherent latent physiological state, together with its event-conditioned transition dynamics, from heterogeneous and partially observed data. Several complementary learning paradigms may contribute at different stages of PWM development, including model-based reinforcement learning, observation-level generative modelling, latent-dynamics learning, and JEPA-style predictive learning [26, 38, 59, 60]. Observation-level generative world models generate future trajectories in raw observation space, including medical images, ECG, PPG, sleep signals, or multimodal sequences [38,65,66]. They can learn complex temporal patterns, but these patterns do not necessarily correspond to intervention-relevant physiological states. Physiological data contain device noise, sampling variation, motion artefacts, and fluctuations unrelated to the individual’s underlying physiological state. A reconstruction-heavy objective may therefore emphasize local morphology rather than underlying physiological dynamics [59, 60]. For this reason, we argue that latent-dynamics learning and JEPA-style prediction provide promising starting points for early-stage PWM development, potentially in combination with observation-level reconstruction objectives. These methods encode high-dimensional observations into latent states and predict future representations rather than every raw detail [59,60]. The model should preserve information relevant to physiological state and intervention, not every low-level fluctuation. At later capability levels, model-based reinforcement learning provides an important conceptual reference for PWM development. Physiological condition can be represented as a latent state, while observed behaviours and assigned treatments can be represented as actions or interventions. Environmental changes may be represented as WorldEvents or contextual variables, and as Interventions when deliberately manipulated. The model can then learn transitions conditioned on these variables and simulate possible future trajectories [42]. However, health data are predominantly observational and offline. Real-world interaction is costly and ethically constrained; rewards are difficult to define; and the effects of observed actions are often confounded. Model-based reinforcement learning is therefore better viewed as a potential route towards L4 planning and closed-loop control rather than as the sole starting point for early-stage training [26, 42, 67–69]. A practical development path can be organized into three stages. First, multimodal self-supervised learning, including JEPA-style objectives, is used to construct stable latent representations of HumanState. Building on these representations, HumanState Transition Tokens are then used to learn how physiological states evolve in response to specific events, actions, and interventions within defined contexts. Finally, the learned transition dynamics may support intervention comparison and planning, but only when high-confidence transition data, causally informative evidence, calibrated uncertainty, and explicit safety constraints are available. Physiology-informed constraints and prospective validation become increasingly important as the model advances towards counterfactual simulation and closed-loop use [26]. 3.2 Capability Levels of PWMs Having reviewed the learning paradigms, we now characterize PWMs according to their functional capabilities through a four-level ladder, summarized in Figure 3. These levels describe functional capabilities in physiological modelling, rather than degrees of autonomy, clinical readiness, or regulatory maturity. This ladder progresses from L1 state representation to L4 bounded intervention planning; each level builds on the previous one and demands progressively stronger evidence, ranging from unlabelled observational data to causally informative intervention data and prospective validation. 3.2.1 L1: State Representation An L1 PWM learns robust and temporally coherent latent representations of HumanState from multimodal physiological observations and interpretable physiological variables [70,71]. Inputs may include PPG, ECG, CGM, respiration, EDA, activity, and validated sleep-related measures [57,58,70–72]. Time-aligned environmental measurements may be used as contextual inputs, but they remain distinct from the latent physiological state [9, 10]. 6 Figure 3 Capability ladder for Physiological World Models. Model capability increases vertically from L1 state representation through L2 state-transition prediction and L3 simulation and counterfactual rollout to L4 planning and intervention. The scope of temporal reasoning expands horizontally from the current state to single transitions, multi-step trajectories and closed-loop updating. The nested bands indicate the progressively broader functional scope of each capability level. The primary output of L1 is a latent state representation or a set of interpretable state estimates, such as recovery, fatigue, or metabolic stability. L1 can leverage large-scale unlabelled or weakly labelled data, yet multimodal synchronization, sparse annotation, and quality control remain central challenges. Evaluation should therefore emphasize cross-device transfer, generalization across populations, temporal consistency, agreement with established physiological measures, and robustness to noise or missing modalities [73–75]. 3.2.2 L2: State-Transition Prediction An L2 PWM learns an event- or action-conditioned transition distribution froms t tos t+∆t . Unlike forecasting based only on prior observations, it predicts how the latent physiological state evolves in response to a specified event, action, or intervention, conditional on the individual’s current state and context. L2 therefore relies on HumanState Transition Tokens to learn conditional transition dynamics. An L2 model may predict latent-state transitions following exercise, a meal, sleep restriction, or acute stress, together with observable outcomes such as heart-rate recovery or postprandial glucose response. It requires clearly defined baseline and response windows, reliable timestamps and event, action, or intervention annotations, sufficient temporal coverage, and contextual information [76, 77]. Without event or behaviour labels, it may collapse into ordinary forecasting. Because L2 predicts conditional trajectories rather than isolated point estimates, evaluation should extend beyond pointwise error to assess trajectory shape, response delay, peak response, recovery speed, within-person consistency, and generalization across contexts [77]. 3.2.3 L3: Simulation and Counterfactual Rollout L3 generates multiple possible future trajectories under alternative events, actions, or interventions. With observational data alone, these outputs should be interpreted as scenario-conditioned rather than counterfactual rollouts. Interpreting these outputs as counterfactual rollouts or estimates of causal effects requires interventional or otherwise causally informative evidence, together with explicit assumptions regarding potential unmeasured confounding, temporal carryover, adherence, and overlap between comparison groups [78–81]. The central challenge is therefore to determine whether the available evidence supports the proposed alternative-trajectory comparison. L3 requires high-confidence HumanState Transition Tokens with reliable event timing, complete baseline and response windows, relevant contextual information, and adequate coverage across individuals and settings. Evaluation should 7 assess factual trajectory accuracy, agreement with held-out intervention responses or other empirical validation targets, intervention-ranking accuracy, uncertainty calibration, and recognition of queries outside the model’s valid range [82–86]. At deployment, population-level distribution-shift monitoring should complement these query-level checks [87]. 3.2.4 L4: Planning and Intervention L4 extends simulation to bounded planning. A planning layer built around the PWM compares candidate interventions under explicit objectives, costs, contraindications, and safety constraints. It ranks or recommends options for human review and updates these estimates as new observations become available [26,88,89]. L4 requires defined target populations, calibrated uncertainty, interpretable ranking logic, prospective validation, intervention feedback, and evaluation of failure cases and rare adverse responses [46,47,90,91]. Post-deployment distribution-shift monitoring should complement, but not substitute for, these prospective evaluations [87]. Lower-risk domains such as sleep scheduling, exercise recovery planning, and environmental management may provide suitable early testbeds. Higher-risk uses, including medication adjustment or treatment selection, require clinical oversight, ethical review, and clearly defined responsibility boundaries; in these settings, intervention selection and execution should not be autonomous [92]. Taken together, L1–L4 describe progressively stronger capabilities for modelling physiological state and its transitions. L1 constructs the state representation. L2 predicts event-, action-, or intervention-conditioned transitions within defined contexts. L3 compares alternative future trajectories and, when supported by causally informative evidence, evaluates counterfactual interventions. L4 uses validated transition dynamics to support bounded intervention planning. The associated data and evidence requirements also become progressively more structured: L1 can use large-scale unlabelled or weakly labelled recordings, whereas L2 relies on high-confidence HumanState Transition Tokens, L3 additionally requires causally informative intervention evidence, and L4 further requires intervention feedback and prospective safety evidence. 4 Data Protocols for Physiological World Models The PWM framework requires temporally aligned physiological observations and state-transition evidence that existing digital-health infrastructure does not routinely provide. Relevant data may combine continuous sensor streams with event annotations, contextual measurements, intervention assignments, and longitudinal outcomes. The value of such datasets lies not in temporal density alone, but in the temporal alignment that enables characterization of how physiological states change within and across individuals. Together, these limitations constrain PWM development. First, despite the widespread adoption of wearable devices and digital-health platforms, these data remain fragmented across devices and institutional systems [3,93–95]. Hospital information systems, consumer wearables, and smartphone applications use different communication protocols, sampling schemes, and storage formats. Standards such as FHIR, Open mHealth, and the OMOP common data model improve interoperability, but additional specifications are needed to align dense, event-anchored, multimodal time series [96–98]. Figure 4 illustrates why interoperability alone is insufficient for PWM development. In fragmented digital-health records, physiological signals and contextual annotations remain distributed across devices, applications, and clinical systems, often with isolated timestamps and poorly resolved cross-modal relationships. Protocol-based temporal alignment instead maps multimodal observations, events, actions, and interventions onto a shared timeline. This common temporal reference defines pre-event baselines and post-event response windows and links physiological changes to their relevant context and observed outcomes, thereby enabling the construction of comparable HumanState Transition Tokens. A second limitation is that event annotations and intervention records are often sparse, incomplete, or temporally imprecise. A continuous glucose monitor may capture a postprandial glucose excursion with minute-level resolution, while the meal itself may be unlogged or its timing inaccurately self-reported [99]. Ecological momentary assessment can reduce retrospective recall bias and collect contextual information in the moment, but it still relies on participant reporting [100]. Similarly, a wearable ECG device may capture a transient rhythm change without a reliable record of concurrent stressors, activities, or environmental exposures. Such omissions limit the construction of event-anchored transition samples from otherwise extensive physiological recordings. Third, physiological response windows vary across events and individuals. Exercise, meals, and sleep disruption may produce different response magnitudes, delays, and recovery patterns depending on baseline physiology, disease status, and circadian phase [72,101]. Without standardized definitions of pre-event baselines, post-event response windows, outcomes, and data quality, state transitions cannot be compared reliably across individuals, devices, or studies. 8 Figure 4 From fragmented digital-health records to protocol-based temporal alignment. A, Wearable, metabolic, clinical, behavioural, and environmental data are recorded asynchronously across devices and applications. B, First-person vision, audio, physiological signals, glucose, ECG, sleep, environmental context, activity, and clinical records are organized on a shared event-anchored timeline. Temporal alignment links events and interventions to pre-event baselines and subsequent physiological responses, enabling the construction of comparable HumanState Transition Tokens. Finally, existing datasets remain dominated by passively collected, free-living recordings, often from a single modality. These data are valuable for characterizing baseline distributions and learning latent state representations, but they are insufficient to identify causal intervention effects or to support validation of L3–L4 capabilities [3]. Stronger evidence requires structured event annotation, repeated measurements, assigned interventions where ethical and feasible, and prospective observation of subsequent outcomes. These limitations motivate a minimum specification for HumanState Transition Tokens and a four-level framework for data acquisition and validation. The protocol levels provide progressively stronger forms of state-transition evidence, but they do not map one-to-one onto PWM capability levels. 4.1 Minimum Data Protocol To enable HumanState Transition Tokens to serve as standardized and comparable data units across studies, devices, and populations, their constituent fields must be defined with sufficient precision to allow independently collected tokens to be pooled for model training and benchmark evaluation [102,103]. We therefore define nine minimum protocol components, each accompanied by examples, minimum recording requirements, and the principal failure modes that arise when those requirements are not met (Table 1). This Minimum Data Protocol serves two purposes. First, it provides a concrete specification that the data producers, such as researchers, device manufacturers, and clinical programmes, can use to design and evaluate data collection work [103,105]. Second, it defines the minimum quality requirements that determine whether a token is suitable for training or evaluation at a given capability level [102, 105]. 4.2 Data Acquisition and Validation Levels Drawing on fit-for-purpose evaluation principles for connected sensor technologies [102,105] and FAIR data-stewardship principles [103], we propose four PWM-specific levels of data acquisition and validation, designated P1 through P4. These levels differ in the precision of event annotation, the degree of intervention control, the availability of prospective feedback, and the strength of the state-transition evidence they provide. Their relationship to the PWM capability levels is enabling rather than one-to-one: a capability may depend on evidence from multiple protocol levels, and no protocol level alone guarantees that the corresponding model capability has been achieved. 9 Table 1 Minimum data requirements and associated failure modes for HumanState Transition Tokens. Component Examples Required precision Failure modes Time Synchronization All sensor streams aligned to a common UTC-anchored timeline Clock drift and timestamp accuracy documented;synchronization accuracy matched to themodality and expected physiological responsetimescale Temporal misalignment between event markersand physiological response windows; spuriouslead-lag relationships Sensor Modalities PPG, ECG, CGM, actigraphy, respiration, EDA, ambient temperature, light, audio level Device model and firmware version documented;sampling rate and filter settings specified Inability to pool data across sensor generations;unmodelled differences in signal fidelity Event Ontology Noise exposure, meeting (type, intensity), ambient temperature change, light exposurechange, social interaction, air quality Event category drawn from a standardized,hierarchical taxonomy of external events andexposures; event onset or exposure intervalrecorded to the required accuracy Event ambiguity across studies; inability tocompare “commute” responses when commutemode, duration, and peak stressor timing differ;unmeasured external events that explainphysiological fluctuations misattributed tospontaneous variation Action Log Eating (food type and approximate quantity), exercise (type, intensity, duration), medication (drug, dose), caffeine, alcohol, screen use, sleep onset/offset Structured, timestamped log entries; smartphone-based or voice-assisted entry used where feasibleto minimize reliance on retrospective recall Recall bias; incomplete or incorrect actionlabelling that contaminates event-conditionedlearning; unrecorded actions that leavephysiological traces in the data but are absentfrom the log, causing the model to attribute theireffects to other variables Context Variables Age, sex, chronic conditions, medications, fitnesslevel, chronotype, baseline HR/HRV/CGMsetpoints from a pre-specified observation periodappropriate to the modality and task; location,activity state, prior sleep duration and quality,circadian phase, ambient conditions Standardized demographic and clinical coding;location from GPS or beacon at a spatial resolution appropriate to the contextual exposureof interest; sleep duration and quality derivedfrom wearable sleep staging validated againstreference polysomnography [104]; circadianphase estimated from activity and light exposure Baseline physiology recorded during anunrepresentative period (e.g., acute illness, travel);undocumented conditions or unreportedmedications that alter response dynamics; omittedsituational moderators (e.g., location, prior sleepdebt) that explain response heterogeneity acrossotherwise identical tokens Intervention Record Dietary regimen, exercise prescription, sleep schedule adjustment, medication change Intervention type, dose or intensity, adherencemetrics, start and end timestamps Inability to distinguish intervention effects fromspontaneous physiological variation; unmeasurednon-adherence Baseline Observation Window Continuous physiological observations over apre-specified window preceding the focal event,action, or intervention, with a task-specificmissingness threshold Sufficient to compute baseline HR, HRV, glucose,activity level, and stress indices Unstable or unrepresentative baselinecontaminating the estimated state-transitionmagnitude Physiological Response Window Continuous physiological observations over apre-specified window following the focal event,action, or intervention Duration matched to the expected physiologicalresponse timescale of the event category;missingness < 20% Truncated response window; missed delayedeffects; inability to characterize responsedynamics Outcome and DataQuality Quantifiable physiological change (e.g.,postprandial glucose AUC, HRV recovery slope,sleep efficiency) plus documented qualityindicators covering completeness, temporalalignment, sensor quality, annotation reliability,and overall confidence Outcome measure defined a priori for each eventcategory; quality score computed from the fivehigh-confidence token dimensions Uninterpretable or misleading transition samplesentering the training corpus; inflated or unstablemodel performance estimates 10 P1 — Passive Free-Living Recording. At this level, data are collected passively during participants’ normal daily lives without assigned activities or experimental interventions. The primary objective is to characterize baseline physiological distributions and their natural temporal variability. Relevant measurements may include heart rate, HRV, glucose, activity, estimated sleep stages, and their circadian and day-to-day patterns. P1 data are often noisy. Events are often sparsely labelled or inferred from sensor patterns, while contextual information and intervention records may be incomplete. P1 therefore provides the observational foundation for self-supervised or weakly supervised learning of latent HumanState representations, but it does not necessarily produce high-confidence HumanState Transition Tokens. Large-scale cohorts with wearable components, such as the UK Biobank accelerometry study [106] and the All of Us Research Program [107], illustrate important elements of P1 data collection at population scale. P2 — Structured Event and Action Annotation Protocol. At P2, participants remain in free-living settings but follow structured event- and action-logging and measurement procedures. Meals, exercise sessions, sleep timing, and other relevant events are recorded using prespecified fields and reliable timestamps. Examples include meal logs supported by photographs, timestamped exercise sessions with intensity ratings, sleep diaries, and in-the-moment self-documentation or ecological momentary assessment [100]. P2 does not require the assignment of an intervention condition. The objective of P2 is to construct comparable, event-anchored transition samples. P2 therefore can provide structured observational data that support L2 event- and action-conditioned state-transition learning. P3 — Controlled Intervention Studies. At P3, specific behavioural, environmental, or clinical interventions are deliberately assigned, and their physiological effects are measured under controlled or semi-controlled conditions. Examples include randomized crossover studies comparing prescribed exercise with habitual activity, studies comparing controlled dietary conditions, and protocols that systematically advance or delay sleep timing [108]. Randomized crossover and repeated-measures designs can provide stronger evidence for within-person intervention effects, provided that randomization, adherence, period effects, carryover effects, and relevant time-varying confounders are appropriately addressed [109,110]. These protocols can produce high-confidence transition tokens containing intervention assignments, pre-intervention states, contextual variables, post-intervention trajectories, outcomes, and data-quality information. P3 provides an important empirical basis for L3 intervention comparison and counterfactual evaluation. Nevertheless, an outcome observed under a different condition or at another time should not be treated as ground truth for the unobserved individual-level counterfactual. P4 — Adaptive Prospective and Shadow-Mode Validation. P4 targets event–context combinations and physio- logical conditions that remain sparsely represented after P1–P3. At this level, model predictions or intervention rankings are generated prospectively before the corresponding outcomes are observed; in shadow mode, they do not influence participant behaviour or clinical care. Subsequent real-world observations are then used to evaluate predictive accuracy, calibration, failure modes, and robustness under distributional shift. Model uncertainty and observed failures may guide additional data collection or ethically bounded intervention studies, directing empirical data collection towards poorly characterized regions of the state-transition space [111]. Clearly labelled synthetic trajectories may complement P4 by supporting hypothesis generation and the exploration of rare scenarios [112], but cannot replace empirical observations, causal evidence, prospective validation or evidence of intervention safety. P4 supports the validation requirements of L3 and L4 rather than independently enabling L4 planning. Any progression towards L4 additionally requires defined target populations, explicit safety constraints, prospective intervention evidence, failure-case analysis, long-term follow-up, and appropriate human or clinical oversight. These levels are complementary rather than strictly cumulative, and a study may combine elements from several levels. Progress from L1 to L4 therefore requires not only larger datasets, but also increasingly structured, causally informative, and prospectively validated evidence. Standardized and reusable data should follow FAIR data-stewardship princi- ples [103], whereas measurement and clinical validation should follow fit-for-purpose evaluation frameworks [102,105]. Causally informative and prospective evidence additionally requires appropriate intervention designs and prospective clinical evaluation [47,78,79]. Progress in physiological world modelling will depend not only on advances in model architecture but also on rigorous data acquisition, well-designed intervention studies, and prospective validation. 11 Table 2 Benchmark tasks for Physiological World Models. TaskEvaluation objectiveRepresentative metricsCapability and evidence requirements T1: HumanState Representation and Estimation Determine whether the model constructs a physiologically meaningful and transferable HumanState representation Agreement with validated physiological measures; cross-device and cross-population transfer; robustness to noise and missing modalities L1; primarily P1 T2: Multi-horizon State- Transition Forecasting Predict how HumanState evolves across multiple timescales following a specified event, action, or intervention, conditional on the current HumanState and Context Trajectory error; effective forecasting horizon; prediction-interval coverage; calibration L2; primarily P2 T3: Event-Conditioned and Individualized Response Prediction Determine whether the model captures person-specific response patterns and heterogeneity following a given event, action, or intervention Response direction; peak amplitude; time to peak; recovery rate; personalization gain L2; P2–P3 T4: Alternative Intervention Simulation Compare plausible trajectories under alternative interventions Held-out intervention error; intervention-ranking accuracy; cross-condition generalization error; uncertainty calibration L3; primarily P3 T5: Bounded Planning and Steerability Rank intervention options under explicit objectives, costs, contraindications, and safety constraints Prospective outcome improvement; intervention efficiency; ranking accuracy; safety violations L4; P3–P4 T6: Reliability under Distribution Shift Determine whether the model recognizes when its predictions become unreliable Robustness gap; calibration error; uncertainty–error association; abstention performance; subgroup failure analysis Cross-cutting across L1–L4; P1–P4, with external or prospective evaluation under relevant distribution shifts 5 Benchmarking and Applications Widely used clinical time-series and biomedical-signal benchmarks evaluate tasks such as in-hospital mortality prediction, physiological decompensation, length-of-stay forecasting, phenotype classification, diagnosis, and digital-biomarker classification [113,114]. Although these tasks remain valuable, they do not establish whether a model can represent an integrated HumanState, predict event-conditioned transitions, simulate trajectories under alternative interventions, or recognize when its predictions are unreliable. We therefore define six complementary benchmark tasks. Tasks 1–5 evaluate progressively stronger modelling capabili- ties, whereas Task 6 assesses reliability across all capability levels. Together, these tasks connect the PWM capability hierarchy, data protocols, and potential applications. 5.1 Benchmark Framework Because latent HumanState is not directly observable, T1 should combine several forms of evidence rather than rely on a single latent-space metric. These may include agreement with validated measures, downstream transfer, robustness to missing modalities, and generalization across devices and populations [70, 73–75, 115]. T2 evaluates whether the model can forecast HumanState trajectories conditioned on events, actions, or intervention across multiple prediction horizons. T3 provides a stricter test of whether the model captures person-specific response heterogeneity and improves prediction through individualization. Forecasting based only on prior observations may serve as a baseline, but should not be treated as evidence of L2 Physiological World Model capability. T4 requires a clear distinction between scenario simulation and counterfactual interpretation, with the latter requiring data from randomized, crossover, or other causally informative study designs. The results observed under the held-out intervention conditions can provide empirical validation targets but do not constitute ground truth for the unobserved individual-level counterfactual [79, 80]. T5 should initially be evaluated prospectively in shadow mode and under explicit safety constraints [47,116]. Evaluation across clinical contexts should precede broader deployment [45]. T6 applies to every preceding task and should 12 Table 3 Application horizons and indicative benchmark profiles. Application horizon Representative applications Required tasks Capability level Evidence and validation requirements Near-termSleep and fatigue monitoring; exercise recovery forecasting; metabolic stability monitoring; stress-related recovery T1, T2, T3 and T6 Primarily L1–L2P1–P2 data; external validation; personalization tests; calibration and applicability boundaries Mid-termMeal- and caffeine-timing guidance; sleep scheduling; exercise planning; work–rest adjustment T1–T4 and T6 Primarily L2–L3P2–P3 data; repeated or N-of-1 observations; crossover interventions; intervention-ranking validation Long-termChronic disease trajectories; adaptive rehabilitation planning; clinical-trial enrichment; bounded clinical decision support T1–T6 Primarily L3–L4P3–P4 evidence; prospective validation; safety gates; clinical oversight; regulatory and accountability requirements function as a reliability gate for higher-risk applications, with explicit assessment of uncertainty, robustness and harmful distribution shifts [48, 84, 117–119]. 5.2 Application-Specific Benchmark Profiles The six benchmark tasks defined above serve complementary evaluation purposes. Translating task performance into application readiness requires a principled mapping between each use case and the appropriate combination of tasks, capability levels, and evidence standards. Table 3 summarizes this mapping for representative applications across near-, mid-, and long-term horizons. These horizons indicate an expected progression in capability, evidence, and governance requirements rather than fixed deployment timelines. No single benchmark task or capability level is sufficient to establish application readiness. For example, personalized recovery management requires reliable state representation, multi-horizon prediction, individualized response modelling, and uncertainty calibration. Behavioural intervention design additionally requires credible comparison of alternative trajectories. Higher-risk clinical applications require the complete benchmark profile, including prospective planning evaluation and reliability under distribution shift. 5.2.1 Near-term: Personalized Health and Recovery Near-term PWM applications are likely to focus on lower-risk personal health and recovery management. By integrating sleep-related measures, HRV, activity load, body-temperature rhythms, dietary events, and environmental context, a PWM could estimate recovery trajectories following sleep restriction, stress exposure, exercise, or metabolic perturbation [72, 120–125]. Potential users include athletes, shift workers, healthcare workers, and other populations exposed to demanding or irregular schedules. These applications should provide probabilistic estimates with explicit uncertainty and applicability boundaries rather than deterministic behavioural instructions. 5.2.2 Mid-term: Behavioural Intervention Design Mid-term applications may focus on the comparison of behavioural interventions, including changes in meal timing, caffeine use, sleep environment, exercise scheduling, and work–rest patterns [72,122,126,127]. The aim is not to identify a universally optimal behaviour, but to compare plausible physiological trajectories under alternative choices while accounting for the individual’s current state, baseline, history, and context. In the absence of causally informative intervention data, these comparisons should be described as scenario simulations; causal claims require data from randomized N-of-1 or crossover studies, or other study designs that support causal identification [128, 129]. 13 5.2.3 Long-term: Higher-risk Clinical Decision Support Longer-term applications may include chronic disease trajectory modelling, adaptive rehabilitation planning, metabolic disease management, treatment-toxicity monitoring, clinical-trial enrichment, and in silico patient-response simula- tion [43,44,130–134]. In these settings, PWMs should support trajectory exploration, hypothesis generation, and clinician-in-the-loop decision-making rather than autonomous treatment selection. Risk estimates may be derived from predicted trajectories, but the defining capability remains the modelling and comparison of physiological state transitions. Simulated trajectories should not be treated as replacements for clinical trials or as sufficient evidence for independent treatment decisions. Higher-risk applications require prospective validation, safety-gated evaluation, calibration assessment, failure-case analysis, and clearly defined target populations and accountability boundaries [46,47,90,135]. Applications involving treatment selection, medication adjustment, or intervention planning additionally require regulatory oversight and continuing clinical supervision. Across these horizons, evidence requirements progress from reliable state representation and transition prediction to causally informative intervention evaluation, bounded planning, prospective validation, and robust recognition of model limitations. 6 Challenges, Open Questions, and Scope Boundaries The central challenge for PWMs is not simply increasing data volume. It is establishing reliable evidence for physiological state representation, state-transition prediction, and intervention comparison from heterogeneous, biased, and incomplete real-world records. Six areas require particular attention. 6.1 Data Heterogeneity and Evidence Quality The data bottleneck differs across capability levels. L1 requires large-scale physiological observations that are sufficiently harmonized for representation learning, whereas L2–L4 increasingly depend on well-aligned and verifiable HumanState Transition Tokens. Wearables, CGMs, smartphones, sleep-monitoring devices, and clinical records differ in sampling frequency, device algorithms, missingness patterns, and annotation quality. Motion artefacts, non-wear periods, event ambiguity, and population bias can make measurements of the same physiological variable difficult to compare across devices and populations [136–140]. Future data infrastructure should therefore prioritize interoperability, provenance, temporal alignment, and the quality of state-transition evidence rather than recording duration alone. Device characteristics, preprocessing procedures, event definitions, missingness patterns, and data-quality indicators should be documented in forms that allow observations and transition tokens to be compared across studies and populations. 6.2 Temporal Confounding and Causality Real-world events rarely occur in isolation. Sleep loss, stress, meals, caffeine, exercise, medication use, and environmental exposure often co-occur and affect physiological state with different delays. From observational trajectories alone, it is difficult to distinguish genuine intervention effects from spontaneous recovery and confounding by co-occurring exposures. Repeated within-person measurements and randomized N-of-1 or crossover designs can strengthen evidence about intervention effects [78]. Negative controls can help detect residual confounding and bias in observational analy- ses [141]. Causal sensitivity analyses should assess how conclusions change under plausible levels of unmeasured confounding [142]. Because controlled study conditions differ from everyday life, evidence from these designs should be complemented by evaluation in free-living settings, where prediction reliability can be monitored across devices, populations, and timescales. 6.3 Personalization and Continual Adaptation An individual’s physiological baseline is not fixed. Ageing, disease progression, training load, medication, seasonality, infection, and changes in daily routine can alter how the same event translates into a physiological response. PWMs therefore require continual calibration, but continual adaptation can also introduce model drift and safety risks [143]. 14 Future systems should separate state updating, individual calibration, model-parameter updating, and the generation of high-risk recommendations. These processes require separate validation and monitoring procedures. Continual adaptation should also include model-version tracking, rollback mechanisms, and checks for catastrophic forgetting, calibration drift, and performance degradation across population subgroups [87, 116, 119, 144, 145]. 6.4 Safety, Interpretability, and Trust PWM predictions and model-supported recommendations should not be presented as certain or definitive. They should be accompanied by appropriately calibrated uncertainty estimates, applicability boundaries, and clearly defined abstention or escalation criteria. Users need to understand the basis of lower-risk predictions, while clinicians require sufficient information to accept, reject, or modify model-supported recommendations. Interpretability analyses should relate latent representations of HumanState to validated physiological measures and clinically meaningful outcomes. Feature attribution alone should not be interpreted as evidence of a causal physiological mechanism. Explanations should also remain sufficiently stable across small changes in input, device, and model version to support meaningful review [146–149]. In medical settings, clinician-in-the-loop oversight, prospective validation, transparent reporting, and predefined respon- sibility boundaries are prerequisites for deployment [46,47,90]. Safety governance should address overconfidence, rare adverse responses, distributional shift, inappropriate extrapolation, and failures that may not be captured by average predictive performance. 6.5 Privacy, Governance, and Equity PWMs may combine physiological signals with behaviour, location, sleep, diet, environmental exposure, and, where explicitly included, audio- or image-derived contextual information. Combining these sources can reveal sensitive patterns that would not be apparent in any single data stream, creating privacy and governance risks beyond those associated with any individual data source considered in isolation [150]. Privacy-by-design may include federated or local learning combined with encryption, differential privacy, access control, and traceability [151]. Governance frameworks should additionally specify data-minimization requirements and auditable data-use policies, and should record data provenance, permitted uses, device characteristics, processing histories, access records, and whether a transition sample was empirically observed or synthetically generated. For synthetic data, provenance records should include the generating model, model version, conditioning variables, and intended use. They should also indicate whether a sample was used for training or evaluation. Synthetic samples should remain distinguishable from empirical observations throughout the model lifecycle and should not be treated as intervention or safety evidence. Equity is also essential. PWM development should not privilege users of high-end devices or populations already well represented in digital health datasets. Validation across devices, ages, socioeconomic groups, care settings, and, where relevant to optical sensing, skin tones should form part of the core evaluation framework [152–154]. Performance differences should be reported alongside aggregate results rather than obscured by population-level averages. 6.6 Scope Boundaries and Relation to Digital Twins Digital twins and PWMs are overlapping but distinct. Digital twins generally emphasize an individual-specific virtual representation, continuous state synchronization, and closed-loop interaction with the represented system. PWMs emphasize learnable state-transition dynamics, conditional simulation, and intervention comparison [43,44]. It is therefore too strong to describe digital twins simply as a subset of PWMs or to assume that a personalized PWM automatically constitutes a complete digital twin. A mature PWM could serve as the core dynamics module of some human digital twins. A complete digital twin would additionally require reliable identity mapping, continuous data synchronization, interface design, governance, workflow integration, and clearly defined accountability boundaries. Early PWMs need not be developed as full digital twins. They may begin with lower-risk transition-modelling tasks such as sleep recovery, metabolic stability, exercise recovery, and rehabilitation trajectories. Integration into more complete personalized simulation or decision-support systems should occur only after the relevant capabilities have been prospectively validated. 15 Taken together, these challenges indicate that progress from state representation to intervention planning should be gated not only by model performance, but also by stronger empirical evidence, prospective validation, transparent governance, and clearly defined limits of use. 7 Outlook The near-term value of Physiological World Models does not lie in simulating the human body in its entirety. That goal remains distant and may be unnecessary for macro-level physiological applications. Instead, PWMs can provide an event-conditioned framework that organizes heterogeneous physiological and contextual data around state transitions rather than according to particular devices or institutional sources. Realizing this potential requires a staged roadmap across models, data, and sensing. At the model level, development should proceed from robust multimodal state representation to event-conditioned transition prediction, simulation of alternative trajectories, and bounded planning. Each advance should be gated by calibrated uncertainty, clearly defined applicability limits, and progressively stronger empirical validation. In parallel, longitudinal and event-anchored datasets must span diverse populations, devices, event types, interventions, contexts, and timescales. Standardized protocols, improved sensing accuracy and long-term wearability, cross-device interoperability, and prospective shadow-mode evaluation will be needed to establish whether transition predictions remain reliable across devices and populations. Together, these advances could form a shared infrastructure for computational physiology, with PWMs serving as a unifying computational layer that connects physiological observation, state-transition modelling, and evidence-based evaluation. Such an infrastructure could move health AI beyond describing current states towards estimating conditional future trajectories and, where supported by causally informative evidence, comparing alternative courses of action. References [1] Kennedy, H. L. The evolution of ambulatory ECG monitoring. Prog Cardiovasc Dis 56, 127–132 (2013). [2]Roos, L. G. & Slavich, G. M. Wearable technologies for health research: opportunities, limitations, and practical and conceptual considerations. Brain Behav Immun 113, 444–452 (2023). [3]Hughes, A. M., Taylor, D. J., Morris, P. D. & Brittain, E. L. Wearable devices and cardiovascular health: revolutionizing remote monitoring and disease prevention. Eur Heart J 47, 2130–2145 (2026). [4]Hicks, J. L. et al. Leveraging mobile technology for public health promotion: a multidisciplinary perspective. Annu Rev Public Health 44, 131–150 (2023). [5]Steckhan, N., Broghammer, F. & Powell, D. Sensor wide association studies in digital medicine. npj Digit Med 9, 408 (2026). [6]Li, X. et al. Digital health: tracking physiomes and activity using wearable biosensors reveals useful health-related information. PLoS Biol 15, e2001402 (2017). [7] Battelino, T. et al. Clinical targets for continuous glucose monitoring data interpretation: recommendations from the international consensus on time in range. Diabetes Care 42, 1593–1603 (2019). [8]Steinhubl, S. R. et al. Effect of a home-based wearable continuous ECG monitoring patch on detection of undiagnosed atrial fibrillation: the mSToPS randomized clinical trial. JAMA 320, 146–155 (2018). [9] Torous, J., Onnela, J. P. & Keshavan, M. New dimensions and new tools to realize the potential of RDoC: digital phenotyping via smartphones and connected devices. Transl Psychiatry 7, e1053 (2017). [10] Smets, E. et al. Large-scale wearable data reveal digital phenotypes for daily-life stress detection. npj Digit Med 1, 67 (2018). [11]Booth, B. M. et al. Multimodal human and environmental sensing for longitudinal behavioral studies in naturalistic settings: framework for sensor selection, deployment, and management. J Med Internet Res 21, e12832 (2019). [12] Nemati, E., Batteate, C. & Jerrett, M. Opportunistic environmental sensing with smartphones: a critical review of current literature and applications. Curr Environ Health Rep 4, 306–318 (2017). [13]Coorey, G. et al. The health digital twin to tackle cardiovascular disease—a review of an emerging interdisciplinary field. npj Digit Med 5, 126 (2022). [14] Katsoulakis, E. et al. Digital twins for health: a scoping review. npj Digit Med 7, 77 (2024). [15]Fernandes Prabhu, D., Gurupur, V., Stone, A. & Trader, E. Integrating artificial intelligence, electronic health records, and wearables for predictive, patient-centered decision support in healthcare. Healthcare 13, 2753 (2025). [16] Lourenço Santos, R. & Cruz-Correia, R. J. An HL7 FHIR® IG for lifestyle medicine in learning health systems: Multi-vendor wearable interoperability with documented terminology gaps. Int J Med Inform 217, 106465 (2026). [17] Schüssler-Fiorenza Rose, S. M. et al. A longitudinal big data approach for precision health. Nat Med 25, 792–804 (2019). [18] Acosta, J. N., Falcone, G. J., Rajpurkar, P. & Topol, E. J. Multimodal biomedical AI. Nat Med 28, 1773–1784 (2022). [19] Kline, A. et al. Multimodal machine learning in precision health: a scoping review. npj Digit Med 5, 171 (2022). [20] Topol, E. J. High-performance medicine: the convergence of human and artificial intelligence. Nat Med 25, 44–56 (2019). 16 [21] Esteva, A. et al. A guide to deep learning in healthcare. Nat Med 25, 24–29 (2019). [22] Beam, A. L. & Kohane, I. S. Big data and machine learning in health care. JAMA 319, 1317–1318 (2018). [23] Abràmoff, M. D., Lavin, P. T., Birch, M., Shah, N. & Folk, J. C. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. npj Digit Med 1, 39 (2018). [24] Yu, K.-H., Beam, A. L. & Kohane, I. S. Artificial intelligence in healthcare. Nat Biomed Eng 2, 719–731 (2018). [25] Hafner, D. et al. Learning latent dynamics for planning from pixels. In: Proceedings of the 36th International Conference on Machine Learning (2019). PMLR, vol. 97, p. 2555–2565. [26] Yu, C., Liu, J., Nemati, S. & Yin, G. Reinforcement learning in healthcare: a survey. ACM Comput Surv 55, 1–36 (2023). [27] Hansen, N. A., Su, H. & Wang, X. Temporal difference learning for model predictive control. In: Proceedings of the 39th International Conference on Machine Learning (2022). PMLR, vol. 162, p. 8387–8406. [28]Dawid, A. & LeCun, Y. Introduction to latent variable energy-based models: a path toward autonomous machine intelligence. J Stat Mech 2024, 104011 (2024). [29] Schrittwieser, J. et al. Mastering Atari, Go, chess and shogi by planning with a learned model. Nature 588, 604–609 (2020). [30] Matsuo, Y. et al. Deep learning, reinforcement learning, and world models. Neural Netw 152, 267–275 (2022). [31] Ding, J. et al. Understanding world or predicting future? A comprehensive survey of world models. ACM Comput Surv 58, 1–38 (2026). [32] Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature 616, 259–265 (2023). [33] Abbaspourazad, S. et al. Large-scale training of foundation models for wearable biosignals. In: International Conference on Learning Representations (ICLR) (2024). [34]Rasmy, L., Xiang, Y., Xie, Z., Tao, C. & Zhi, D. Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. npj Digit Med 4, 86 (2021). [35]Gadd, C. et al. SurvivEHR: a competing risks, time-to-event foundation model for multiple long-term conditions from primary care electronic health records. npj Digit Med 9, 546 (2026). [36]Qiu, J., Hu, Y., Li, L. et al. Deep representation learning for clustering longitudinal survival data from electronic health records. Nat Commun 16, 2534 (2025). [37] Ding, S. et al. Integrated sensing networks for One Health monitoring. Nat Sensors 1, 485–493 (2026). [38]Yang, Y., Wang, Z. Y., Liu, Q. et al. Medical World Model. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2025). Pp. 8319–8329. [39]Zhang, H. et al. Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells (2026). arXiv:2603.25240. [40] Zhang, O. et al. ODesign: a world model for biomolecular interaction design (2025). arXiv:2510.22304. [41]Chen, Z. et al.ECG-WM: A Physiology-Informed ECG World Model for Clinical Intervention Simulation (2026). arXiv:2605.17580. [42]Xu, Q., Habib, G., Wu, F., Perera, D. & Feng, M. medDreamer: Model-Based Reinforcement Learning with Latent Imagination on Complex EHRs for Clinical Decision Support. In: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (2026). Pp. 1693–1704. [43]Tudor, B. H. et al. A scoping review of human digital twins in healthcare applications and usage patterns. npj Digit Med 8, 587 (2025). [44] De Domenico, M. et al. Challenges and opportunities for digital twins in precision medicine from a complex systems perspective. npj Digit Med 8, 37 (2025). [45] Li, M. M., Reis, B. Y., Rodman, A. et al. Scaling medical AI across clinical contexts. Nat Med 32, 439–448 (2026). [46]Liu, X., Cruz Rivera, S., Moher, D., Calvert, M. J. & Denniston, A. K. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med 26, 1364–1374 (2020). [47] Vasey, B. et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med 28, 924–933 (2022). [48]Ovadia, Y. et al. Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift. In: Advances in Neural Information Processing Systems 32 (NeurIPS) (2019). Pp. 14003–14014. [49]Hafner, D., Pasukonis, J., Ba, J. & Lillicrap, T. Mastering diverse control tasks through world models. Nature 640, 647–653 (2025). [50] Shaffer, F. & Ginsberg, J. P. An overview of heart rate variability metrics and norms. Front Public Health 5, 258 (2017). [51] Krause, A. J. et al. The sleep-deprived human brain. Nat Rev Neurosci 18, 404–418 (2017). [52] Chellappa, S. L., Vujovic, N., Williams, J. S. & Scheer, F. A. J. L. Impact of circadian disruption on cardiovascular function and disease. Trends Endocrinol Metab 30, 767–779 (2019). [53] Borbély, A. A. The two-process model of sleep regulation: beginnings and outlook. J Sleep Res 31, e13598 (2022). [54] Cella, D. et al. The Patient-Reported Outcomes Measurement Information System (PROMIS) developed and tested its first wave of adult self-reported health outcome item banks: 2005-2008. J Clin Epidemiol 63, 1179–1194 (2010). [55] Zhang, W., Bi, S. & Luo, L. The impact of long-term exercise intervention on heart rate variability indices: a systematic meta-analysis. Front Cardiovasc Med 12, 1364905 (2025). 17 [56]Cole, C. R., Blackstone, E. H., Pashkow, F. J., Snader, C. E. & Lauer, M. S. Heart-rate recovery immediately after exercise as a predictor of mortality. N Engl J Med 341, 1351–1357 (1999). [57] Thapa, R. et al. A multimodal sleep foundation model for disease prediction. Nat Med 32, 752–762 (2026). [58]Luo, Y., Chen, Y., Salekin, A. & Rahman, T. Toward foundation model for multivariate wearable sensing of physiological signals. ACM Trans Comput Healthc 7, 1–43 (2026). [59]Seo, Y. et al. Masked World Models for Visual Control. In: Proceedings of The 6th Conference on Robot Learning (2023). PMLR, vol. 205, p. 1332–1344. [60] Assran, M. et al. Self-Supervised Learning From Images With a Joint-Embedding Predictive Architecture. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023). Pp. 15619–15629. [61] McDermott, M. B. A., Nestor, B., Argaw, P. & Kohane, I. S. Event Stream GPT: a data pre-processing and modeling library for generative, pre-trained transformers over continuous-time sequences of complex events. In: Advances in Neural Information Processing Systems 36 (NeurIPS) (2023). Pp. 24322–24334. [62]Wang, Y. et al. Deep time series models: a comprehensive survey and benchmark. IEEE Trans Pattern Anal Mach Intell 1–20 (2026). [63]Wen, Q. et al. Transformers in time series: a survey. In: Proceedings of the 32nd International Joint Conference on Artificial Intelligence (IJCAI) (2023). Pp. 6778–6786. [64] Kim, J., Kim, H., Kim, H. G., Lee, D. & Yoon, S. A comprehensive survey of deep learning for time series forecasting: architectural diversity and open challenges. Artif Intell Rev 58, 216 (2025). [65]Neifar, N., Mdhaffar, A., Ben-Hamadou, A. & Jmaiel, M. Deep generative models for physiological signals: a systematic literature review. Artif Intell Med 165, 103127 (2025). [66]Rasul, K., Seward, C., Schuster, I. & Vollgraf, R. Autoregressive Denoising Diffusion Models for Multivariate Probabilistic Time Series Forecasting. In: Proceedings of the 38th International Conference on Machine Learning (2021). PMLR, vol. 139, p. 8857–8868. [67] Gottesman, O. et al. Guidelines for reinforcement learning in healthcare. Nat Med 25, 16–18 (2019). [68]Jayaraman, P., Desman, J., Sabounchi, M., Nadkarni, G. N. & Sakhuja, A. A primer on reinforcement learning in medicine for clinicians. npj Digit Med 7, 337 (2024). [69]Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C. & Faisal, A. A. The Artificial Intelligence Clinician learns optimal treatment strategies for sepsis in intensive care. Nat Med 24, 1716–1720 (2018). [70]Gu, X. et al. Cardiac health assessment across scenarios and devices using a multimodal foundation model pretrained on data from 1.7 million individuals. Nat Mach Intell 8, 220–233 (2026). [71] Yuan, W., Jin, Z., Wang, Y. et al. sleep2vec: Unified Cross-Modal Alignment for Heterogeneous Nocturnal Biosignals. In: International Conference on Learning Representations (ICLR) (2026). [72] Berry, S. E. et al. Human postprandial responses to food and potential for precision nutrition. Nat Med 26, 964–973 (2020). [73]Shen, Q., Xin, J., Dai, B., Zhang, S. & Wang, Z. Robust Sleep Staging over Incomplete Multimodal Physiological Signals via Contrastive Imagination. In: Advances in Neural Information Processing Systems 37 (NeurIPS) (2024). Pp. 112025–112049. [74] Perslev, M. et al. U-Sleep: resilient high-frequency sleep staging. npj Digit Med 4, 72 (2021). [75]Ong Ly, C. et al. Shortcut learning in medical AI hinders generalization: method for estimating AI model generalization without external data. npj Digit Med 7, 124 (2024). [76]Shen, X. et al. Multi-omics microsampling for the profiling of lifestyle-associated changes in health. Nat Biomed Eng 8, 11–29 (2024). [77]Collins, G. S. et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 385, e078378 (2024). [78] Vohra, S. et al. CONSORT extension for reporting N-of-1 trials (CENT) 2015 statement. BMJ 350, h1738 (2015). [79]Keogh, R. H. & Van Geloven, N. Prediction under interventions: evaluation of counterfactual performance using longitudinal observational data. Epidemiology 35, 329–339 (2024). [80] Vollenweider, M. et al. Learning personalized treatment decisions in precision medicine: disentangling treatment assignment bias in counterfactual outcome prediction and biomarker identification. In: Proceedings of the 4th Machine Learning for Health Symposium (ML4H) (2025). PMLR, vol. 259, p. 991–1013. [81]Shalit, U., Johansson, F. D. & Sontag, D. Estimating individual treatment effect: generalization bounds and algorithms. In: Proceedings of the 34th International Conference on Machine Learning (2017). PMLR, vol. 70, p. 3076–3085. [82] Gneiting, T. & Raftery, A. E. Strictly proper scoring rules, prediction, and estimation. J Am Stat Assoc 102, 359–378 (2007). [83] Cheon, J. & Paik, S. B. Brain-inspired warm-up training with random noise for uncertainty calibration. Nat Mach Intell 8, 602–613 (2026). [84] Chakraborti, T. et al. Personalized uncertainty quantification in artificial intelligence. Nat Mach Intell 7, 522–530 (2025). [85] Guo, C., Pleiss, G., Sun, Y. & Weinberger, K. Q. On calibration of modern neural networks. In: Proceedings of the 34th International Conference on Machine Learning (2017). PMLR, vol. 70, p. 1321–1330. [86] Angelopoulos, A. N. & Bates, S. A gentle introduction to conformal prediction and distribution-free uncertainty quantification. Found Trends Mach Learn 16, 494–591 (2023). 18 [87]Koch, L. M., Baumgartner, C. F. & Berens, P. Distribution shift detection for the postmarket surveillance of medical AI algorithms: a retrospective simulation study. npj Digit Med 7, 120 (2024). [88] Achiam, J., Held, D., Tamar, A. & Abbeel, P. Constrained policy optimization. In: Proceedings of the 34th International Conference on Machine Learning (2017). PMLR, vol. 70, p. 22–31. [89]Thomas, P., Theocharous, G. & Ghavamzadeh, M. High confidence policy improvement. In: Proceedings of the 32nd International Conference on Machine Learning (2015). PMLR, vol. 37, p. 2380–2388. [90] Sel, K. et al. Survey and perspective on verification, validation, and uncertainty quantification of digital twins for precision medicine. npj Digit Med 8, 40 (2025). [91] Badani, A. et al. AI and innovation in clinical trials. npj Digit Med 8, 683 (2025). [92]Sokol, K., Fackler, J. & Vogt, J. E. Artificial intelligence should genuinely support clinical reasoning and decision making to bridge the translational gap. npj Digit Med 8, 345 (2025). [93] Liu, F. & Panagiotakos, D. Real-world data: a brief review of the methods, applications, challenges and opportunities. BMC Med Res Methodol 22, 287 (2022). [94] Coravos, A. et al. Digital medicine: a primer on measurement. Digit Biomark 3, 31–71 (2019). [95] Migueles, J. H. et al. Accelerometer data collection and processing criteria to assess physical activity and other outcomes: a systematic review and practical considerations. Sports Med 47, 1821–1845 (2017). [96]Li, E., Clarke, J., Ashrafian, H., Darzi, A. & Neves, A. L. The impact of electronic health record interoperability on safety and quality of care in high-income countries: systematic review. J Med Internet Res 24, e38144 (2022). [97]Falkenhein, I. et al. Wearable Device Health Data Mapping to Open mHealth and FHIR Data Formats. Stud Health Technol Inform 305, 341–344 (2023). [98]Papez, V. et al. Transforming and evaluating electronic health record disease phenotyping algorithms using the OMOP common data model: a case study in heart failure. JAMIA Open 4, ooab001 (2021). [99]Popp, C. J. et al. Objective determination of eating occasion timing: combining self-report, wrist motion, and continuous glucose monitoring to detect eating occasions in adults with prediabetes and obesity. J Diabetes Sci Technol 18, 266–272 (2024). [100] Shiffman, S., Stone, A. A. & Hufford, M. R. Ecological momentary assessment. Annu Rev Clin Psychol 4, 1–32 (2008). [101]Noone, J., Mucinski, J. M., DeLany, J. P., Sparks, L. M. & Goodpaster, B. H. Understanding the variation in exercise responses to guide personalized physical activity prescriptions. Cell Metab 36, 702–724 (2024). [102] Goldsack, J. C. et al. Verification, analytical validation, and clinical validation (V3): the foundation of determining fit-for- purpose for Biometric Monitoring Technologies (BioMeTs). npj Digit Med 3, 55 (2020). [103]Wilkinson, M. D. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 3, 160018 (2016). [104]Chinoy, E. D. et al. Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep 44, zsaa291 (2021). [105]Coravos, A. et al. Modernizing and designing evaluation frameworks for connected sensor technologies in medicine. npj Digit Med 3, 37 (2020). [106]Doherty, A. et al. Large scale population assessment of physical activity using wrist worn accelerometers: the UK Biobank Study. PLoS One 12, e0169649 (2017). [107] Denny, J. C. et al. The “All of Us” Research Program. N Engl J Med 381, 668–676 (2019). [108]Hall, K. D. et al. Ultra-processed diets cause excess calorie intake and weight gain: an inpatient randomized controlled trial of ad libitum food intake. Cell Metab 30, 67–77.e3 (2019). [109]Schork, N. J., Beaulieu-Jones, B., Liang, W. S., Smalley, S. & Goetz, L. H. Exploring human biology with N-of-1 clinical trials. Camb Prisms Precis Med 1, e12 (2023). [110] Wang, P. et al. N-of-1 medicine. Singapore Med J 65, 167–175 (2024). [111]Gal, Y., Islam, R. & Ghahramani, Z. Deep Bayesian active learning with image data. In: Proceedings of the 34th International Conference on Machine Learning (2017). PMLR, vol. 70, p. 1183–1192. [112] Chen, R. J., Lu, M. Y., Chen, T. Y., Williamson, D. F. K. & Mahmood, F. Synthetic data in machine learning for medicine and healthcare. Nat Biomed Eng 5, 493–497 (2021). [113]Harutyunyan, H., Khachatrian, H., Kale, D. C., Ver Steeg, G. & Galstyan, A. Multitask learning and benchmarking with clinical time series data. Sci Data 6, 96 (2019). [114]Wang, W. K. et al. A systematic review of time series classification techniques used in biomedical applications. Sensors 22, 8016 (2022). [115] Radhakrishnan, A. et al. Cross-modal autoencoder framework learns holistic representations of cardiovascular state. Nat Commun 14, 2436 (2023). [116]Rosenthal, J. T., Beecy, A. & Sabuncu, M. R. Rethinking clinical trials for medical AI with dynamic deployments of adaptive systems. npj Digit Med 8, 252 (2025). [117]Kompa, B., Snoek, J. & Beam, A. L. Second opinion needed: communicating uncertainty in medical machine learning. npj Digit Med 4, 4 (2021). [118]Balendran, A., Beji, C., Bouvier, F. et al. A scoping review of robustness concepts for machine learning in healthcare. npj Digit Med 8, 38 (2025). 19 [119]Subasri, V., Krishnan, A., Kore, A. et al. Detecting and remediating harmful data shifts for the responsible deployment of clinical AI models. JAMA Netw Open 8, e2513685 (2025). [120] Perez-Pozuelo, I., Zhai, B., Palotti, J. et al. The future of sleep health: a data-driven revolution in sleep science and medicine. npj Digit Med 3, 42 (2020). [121] Thayer, J. F., Åhs, F., Fredrikson, M., Sollers, J. J., I & Wager, T. D. A meta-analysis of heart rate variability and neuroimaging studies: implications for heart rate variability as a marker of stress and health. Neurosci Biobehav Rev 36, 747–756 (2012). [122] Zeevi, D. et al. Personalized nutrition by prediction of glycemic responses. Cell 163, 1079–1094 (2015). [123]Bourdon, P. C. et al. Monitoring athlete training loads: consensus statement. Int J Sports Physiol Perform 12, S2–161–S2–170 (2017). [124] Walsh, N. P. et al. Sleep and the athlete: narrative review and 2021 expert consensus recommendations. Br J Sports Med 55, 356–368 (2021). [125] Roddick, C. M., Seo, Y. S., Barkovich, S. L., Forrester, L. & Chen, F. S. Cardiac vagal recovery following acute psychological stress in human adults: a scoping review. Neurosci Biobehav Rev 176, 106268 (2025). [126]Drake, C., Roehrs, T., Shambroom, J. & Roth, T. Caffeine effects on sleep taken 0, 3, or 6 hours before going to bed. J Clin Sleep Med 9, 1195–1200 (2013). [127] Stothard, E. R. et al. Circadian entrainment to the natural light-dark cycle across seasons and the weekend. Curr Biol 27, 508–513 (2017). [128] Chandereng, T. et al. in Role of digital healthcare approaches in the analysis of personalized (N-of-1) trials (eds Hsueh, P., Wetter, T. & Zhu, X.) Personal Health Informatics 131–146 (Springer, Cham, 2022). [129] Schork, N. J. Personalized medicine: time for one-person trials. Nature 520, 609–611 (2015). [130] Corral-Acero, J. et al. The ’Digital Twin’ to enable the vision of precision cardiology. Eur Heart J 41, 4556–4564 (2020). [131]Venkatesh, K. P., Raza, M. M. & Kvedar, J. C. Health digital twins as tools for precision medicine: considerations for computation, implementation, and regulation. npj Digit Med 5, 150 (2022). [132]Kovatchev, B. P. et al. Human-machine co-adaptation to automated insulin delivery: a randomised clinical trial using digital twin technology. npj Digit Med 8, 253 (2025). [133] Akbarialiabad, H. et al. Enhancing randomized clinical trials with digital twins. npj Syst Biol Appl 11, 110 (2025). [134]Brown, S. A. et al. Six-month randomized, multicenter trial of closed-loop control in type 1 diabetes. N Engl J Med 381, 1707–1717 (2019). [135] Wiens, J. et al. Do no harm: a roadmap for responsible machine learning for health care. Nat Med 25, 1337–1340 (2019). [136] Bent, B., Goldstein, B. A., Kibbe, W. A. & Dunn, J. P. Investigating sources of inaccuracy in wearable optical heart rate sensors. npj Digit Med 3, 18 (2020). [137]Canali, S., Schiaffonati, V. & Aliverti, A. Challenges and recommendations for wearable devices in digital health: data quality, interoperability, health equity, fairness. PLOS Digit Health 1, e0000104 (2022). [138]Chen, J. et al. Barriers to translating continuous monitoring technologies for preventative medicine. Nat Biomed Eng 9, 1797–1815 (2025). [139] Güntner, A. T. et al. Challenges and opportunities of wearable molecular sensors in endocrinology and metabolism. Nat Rev Endocrinol 22, 50–60 (2026). [140]Daniore, P. et al. From wearable sensor data to digital biomarker development: ten lessons learned and a framework proposal. npj Digit Med 7, 161 (2024). [141]Lipsitch, M., Tchetgen Tchetgen, E. & Cohen, T. Negative controls: a tool for detecting confounding and bias in observational studies. Epidemiology 21, 383–388 (2010). [142] Ding, P. & VanderWeele, T. J. Sensitivity analysis without assumptions. Epidemiology 27, 368–377 (2016). [143] Guo, L. L. et al. EHR foundation models improve robustness in the presence of temporal distribution shift. Sci Rep 13, 3767 (2023). [144]Feng, J. et al. Clinical artificial intelligence quality improvement: towards continual monitoring and updating of AI algorithms in healthcare. npj Digit Med 5, 66 (2022). [145]Pesapane, F., Rotili, A., Penco, S., Nicosia, L. & Cassano, E. Evidence over explanations: put medical AI to the test. npj Artif Intell 2, 53 (2026). [146] Adebayo, J. et al. Sanity checks for saliency maps. In: Advances in Neural Information Processing Systems 31 (NeurIPS) (2018). Pp. 9505–9515. [147]Ghorbani, A., Abid, A. & Zou, J. Interpretation of neural networks is fragile. Proc AAAI Conf Artif Intell 33, 3681–3688 (2019). [148]Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell 1, 206–215 (2019). [149] Amann, J. et al. Explainability for artificial intelligence in healthcare: a multidisciplinary perspective. BMC Med Inform Decis Mak 20, 310 (2020). [150] Price, W. N. & Cohen, I. G. Privacy in the age of medical big data. Nat Med 25, 37–43 (2019). [151] Rieke, N. et al. The future of digital health with federated learning. npj Digit Med 3, 119 (2020). [152] Walter, J. R., Xu, S. & Rogers, J. A. From lab to life: how wearable devices can improve health equity. Nat Commun 15, 123 (2024). 20 [153]Nagappan, A., Krasniansky, A. & Knowles, M. Patterns of ownership and usage of wearable devices in the United States, 2020-2022: survey study. J Med Internet Res 26, e56504 (2024). [154] Obermeyer, Z., Powers, B., Vogeli, C. & Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366, 447–453 (2019). 21