Paper deep dive
Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic
Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/23/2026, 1:46:43 AM
Summary
The paper introduces Foresight-England (Foresight-E), a 243-million-parameter generative transformer foundation model trained on de-identified electronic health records (EHRs) of approximately 61 million individuals in England. Developed within the NHS England Secure Data Environment, the model predicts future medical events autoregressively using a zero-shot approach. The study details the methodology for tokenization, architecture, and evaluation, focusing on predicting direct and indirect effects of the COVID-19 pandemic. However, quantitative results are unavailable because NHS England paused data access due to governance concerns raised by medical bodies, though the Information Commissioner's Office concluded no GDPR breach occurred.
Entities (13)
Relation Signals (11)
Foresight-England â architecture â Transformer Decoder
confidence 95% ¡ Foresight-E is a 243-million-parameter transformer decoder
Foresight-England â hostedin â NHS England Secure Data Environment
confidence 95% ¡ Trained from scratch entirely within the NHS England Secure Data Environment
NHS England â pausedaccessfor â Foresight-England
confidence 95% ¡ NHS England has paused access to data for the Foresight-E project
Foresight-England â purpose â COVID-19
confidence 95% ¡ developed as a research pilot strictly for COVID-19 research
Foresight-England â trainedon â Electronic Health Records
confidence 95% ¡ Trained from scratch entirely within the NHS England Secure Data Environment... on de-identified, longitudinal EHRs
Information Commissioner's Office â reviewedcompliancewith â GDPR
confidence 90% ¡ The ICO closed its review of the project and concluded that the projectâs data use was compatible with the purpose for which it was originally shared therefore there was no breach of GDPR.
Foresight-England â usesstandard â OPCS-4
confidence 90% ¡ Our tokenisation scheme retains the clinical granularity of ... OPCS-4
Foresight-England â usesstandard â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research. We evaluated its ability to model the direct and indirect effects of the pandemic. Trained from scratch entirely within the NHS England Secure Data Environment, Foresight-E is a 243-million-parameter transformer decoder. It was trained and evaluated on de-identified, longitudinal EHRs of approximately 61 million individuals, integrating primary/secondary care, death registrations, and COVID-19 data. Training and validation used a 90% subset (54.9 million) spanning November 2018 to December 2022; the remaining 10% (6.1 million) was held out for evaluation. Foresight-E models patient timelines autoregressively, predicting the next medical event given their prior history. At inference, it operates zero-shot, predicting any concept in its ~40,000-code vocabulary without task-specific training. Our tokenisation scheme retains the clinical granularity of ICD-10, OPCS-4, and SNOMED CT codes, jointly representing absolute and relative timing. We designed an evaluation framework for 30-day COVID-19 hospitalisation and mortality, including subgroup analyses by demographic factors and vaccination status. To assess generalisation to unseen future data and the pandemic's indirect effects, we tested the model on medical events from 2023 (beyond its training period), benchmarking against logistic regression and XGBoost. As detailed in the Project Status section, NHS England has paused access to data for the Foresight-E project, meaning quantitative results are currently unavailable. Instead, we share our strategy for tokenisation, architecture, training, inference, and evaluation as a methodological template and case study in the challenges of building population-scale EHR foundation models.
Tags
Links
- Source: https://arxiv.org/abs/2608.16273v1
- Canonical: https://arxiv.org/abs/2608.16273v1
Trouble viewing inline? Open PDF directly â
Full Text
108,412 characters extracted from source content.
Expand or collapse full text
Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic Simon Ellershaw 1, 2 &Christopher Tomlinson* 1, 2, 3, 4 &Zeljko Kraljevic 2 &Spiros Denaxas 1, 4, 5, 6, 7 &Harry Hemingway 1, 4, 7 &Cathie Sudlow 8 &Angela M. Wood 6, 9, 10, 11, 12, 13, 14 &Anoop D. Shah 1, 4, 15 &Richard Dobson 1, 2, 4, 7, 16 &on behalf of the CVD-COVID-UK/COVID-IMPACT Consortium 1 Institute of Health Informatics, University College London, London, UK 2 Department of Biostatistics and Health Informatics, Institute of Psychiatry, Psychology and Neuroscience, Kingâs College London, London, UK 3 Kingâs Institute for Artificial Intelligence, Kingâs College London, London, UK 4 University College London Hospitals National Institute for Health Research Biomedical Research Centre, London, UK 5 Interdisciplinary Transformation University, Linz, Austria 6 British Heart Foundation Data Science Centre, Health Data Research UK, London, UK 7 Health Data Research UK, London, UK 8 Usher Institute, School of Population Health Sciences, The University of Edinburgh 9 British Heart Foundation Cardiovascular Epidemiology Unit, Department of Public Health and Primary Care, University of Cambridge, Cambridge, UK 10 Victor Phillip Dahdaleh Heart and Lung Research Institute, University of Cambridge, Cambridge, UK 11 British Heart Foundation Centre of Research Excellence, University of Cambridge, Cambridge, UK 12 National Institute for Health and Care Research Blood and Transplant Research Unit in Donor Health and Behaviour, University of Cambridge, Cambridge, UK 13 Health Data Research UK Cambridge, Wellcome Genome Campus and University of Cambridge, Cambridge, UK 14 Cambridge Centre of Artificial Intelligence in Medicine, University of Cambridge, Cambridge, UK 15 Department of Clinical Pharmacology, University College London Hospitals NHS Foundation Trust, London, UK 16 National Institute for Health Research Biomedical Research Centre at South London and Maudsley NHS Foundation Trust and Kingâs College London, London, UK Thanks: These authors contributed equally to this work. Thanks: Corresponding author: christopher.tomlinson@ucl.ac.uk Abstract Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19-related research. We evaluated its ability to model the direct and indirect effects of the COVID-19 pandemic. Trained from scratch entirely within the NHS England Secure Data Environment, Foresight-E is a 243-million-parameter transformer decoder. It is trained and evaluated on a national-scale, de-identified, longitudinal EHR dataset of approximately 61 million individuals, integrating primary and secondary care, death registrations, and COVID-19 testing/vaccination datasets. Training and validation were conducted on a random subset of 90% of individuals (54.9 million) for events recorded between 1st November 2018 and 31st December 2022. The remaining 10% of individuals (6.1 million) were held out for evaluation. Foresight-E models patient timelines autoregressively, predicting the next medical event given an individualâs prior history. At inference, it operates zero-shot, generating predictions for any concept in its approximately 40,000-code medical vocabulary without additional task-specific training. Our custom tokenisation scheme retains the recorded clinical granularity of ICD-10, OPCS-4, and SNOMED CT codes, and jointly represents both absolute (e.g., calendar dates) and relative timing (e.g., chronological age). We designed and implemented an evaluation framework spanning 30-day COVID-19 hospitalisation and mortality using Brier scores and area under the receiver operating characteristic (AUROC) and precision-recall (AUPRC) curves. Subgroup analysis of these results by age, ethnicity, sex and COVID vaccination status was also conducted. We also tested Foresight-E on medical events from 2023, extending beyond its 2018â2022 training period, to assess how well it captured the enduring, system-wide indirect effects of the pandemic on future unseen data, simulating a prospective deployment. We benchmarked model performance against logistic regression and XGBoost. As detailed in the Project Status section, NHS England has paused access to data for the Foresight-E project, meaning quantitative results are not currently available. Instead, we share our strategy for tokenisation, model architecture, training, inference, and evaluation, as a methodological template and a case study in the challenges of building population-scale EHR foundation models. 1 Lay summary Predicting who gets ill, when, and with which diseases are vital questions for individuals, doctors, and healthcare systems, such as the National Health Service (NHS). The COVID-19 pandemic urgently highlighted this need. The NHS had to quickly identify people who might become very sick if they caught the virus, so they could be prioritised for vaccinations or treatments. While many predictive tools, including some using artificial intelligence (AI), were developed during the pandemic, they often focused on single, narrow questions, like predicting the risk of death only after a patient was admitted to hospital. Because they couldnât look at the bigger picture, they struggled to capture the wider, long-term impacts of the pandemic on the healthcare system. To address this, we developed a new AI system called Foresight-England (Foresight-E) to predict a patientâs future medical events. Our tool learns from past patient data to estimate what might happen to a patientâs health in the future, and when these events might occur. It works a bit like the predictive text on a mobile phone, which tries to predict the next word in a sentence, but it uses medical codes instead of words. Our goal was to test Foresight-Eâs ability to model changes in peopleâs health during the pandemic. We looked at two main areas: ⢠Direct Effects of COVID-19: We tested how accurately the AI could tell us which patients would be hospitalised or die within 30 days of a positive COVID-19 test. We also checked if the tool could learn and adapt to the shifting nature of the pandemic, such as new viral variants and the vaccine rollout. Importantly, we checked if the AI worked fairly across different demographic groups, looking for evidence of bias in the underlying data that might lead to unequal predictions. ⢠Indirect Effects of the Pandemic: The pandemic caused severe disruptions to routine healthcare, leaving lasting effects. To see if the tool could predict these broader impacts, we trained it on data up to the end of 2022, and then tested it on "unseen" data from 2023. We looked at emergency hospital admissions, overall deaths, and the new onset of over 1,400 different diseases whose diagnosis or treatment might have been delayed or triggered by the pandemic. AI models are only as good as the data they learn from. To make sure the model was fair and represented people from all backgrounds, we trained Foresight-E using de-identified electronic health records from 54.9 million peopleârepresenting almost the entire population of England. We then tested the tool on a separate group of 6.1 million people. This data was securely accessed via the British Heart Foundation Data Science Centreâs CVD-COVID-UK/COVID-IMPACT consortium. The data looked a bit like a computer spreadsheet of medical codes and dates, not written text notes from doctors. It included GP records, hospital visits, COVID-19 testing, vaccinations, and national death registrations. The project was designed so that the data, the AI model, and its predictions were kept entirely within a highly secure digital system called the NHS England Secure Data Environment (SDE). This environment uses strict rules (the "Five Safes" framework) meaning no patient data ever left the NHS. The AI was built from scratch inside this secure system. Our industry partners, Amazon Web Services (AWS) and Databricks, provided computer power and technical help, but they had absolutely no access to the data, the AI model, or control over the research. It is important to note that Foresight-E was developed strictly as a research project for COVID-19 and is not used by the NHS for patient care. NHS England has paused access to the data for this project. Because we cannot access the secure environment, we cannot retrieve the initial results of our study. Therefore, we have provided a transparent overview of how we designed and tested the system, alongside "placeholder" mock-ups of our results tables to demonstrate the work we have already completed, and how we were intending to report it. 2 Project Status In May 2025, the British Medical Association and Royal College of General Practitionersâ Joint GP IT Committee (JGPITC) raised concerns that they were unaware that primary care data collected for COVID-19 research was being used to train an AI model, and queried whether the correct processes, including GDPR principles, had been followed [4]. Following these concerns, NHS England paused access to data for the Foresight-E project whilst a governance review was carried out. This review considered whether the correct approval processes and privacy protections were in place for a project of this nature. The JGPITC additionally wrote to the Information Commissionerâs Office (ICO), the UKâs data protection regulator, asking them to investigate. In April 2026, following a thorough review, the ICO closed its review of the project and concluded that the projectâs data use was compatible with the purpose for which it was originally shared therefore there was no breach of GDPR. The projectâs access to the NHS England SDE remains paused. At the point at which data access was paused, the researchers had trained several iterations of the Foresight-E model (including on the full training dataset), generated multiple sets of predictions for direct and indirect COVID-19 outcomes in the test set, and run initial quantitative evaluations on these predictions, producing aggregated performance metrics, as detailed in this manuscript. However these aggregate predictions had not yet been requested for export from the NHS England Secure Data Environment (SDE) via the approved Safe Output Service, subject to statistical disclosure control and review. This means quantitative results are not available, and this paper instead seeks to provide a transparent overview of the underpinning methodology, evaluation strategy and work undertaken to date on Foresight-E. Placeholder results are presented to provide full transparency over the nature of the evaluation and aggregate data (e.g. tables, figures) that were intended to be exported from the SDE following standard procedure. Whilst the researchers describe the work to the best of their ability, it is important to note that since data access was paused, they have been unable to access not only the underlying data, but their codebase, models, experiment tracking, documentation and results, which are stored inside the SDE in platforms such as Databricks, Gitlab and MLflow. 3 Introduction Accurate risk stratification is an important component of clinical decision-making, enabling individuals, healthcare professionals and health systems to anticipate adverse outcomes, tailor interventions, and inform resource allocation. Traditional clinical risk prediction models typically combine established expert knowledge with rule-based or statistical approaches that harness a limited number of static features, such as demographics, pre-existing conditions, or test results, at a single point in time to predict a single outcome [23, 36]. While effective for narrow, well-understood diseases, these methods face fundamental limitations during a global pandemic caused by a novel pathogen. When COVID-19 emerged, healthcare systems lacked a priori knowledge of interacting risk factors, and the epidemiology shifted rapidly with viral variants and changing public health interventions, such as vaccination rollouts. Traditional approaches struggle to model this complexity and cannot easily scale to capture the multitude of ways a systemic shock like COVID-19 impacts the healthcare system across a vast array of possible clinical outcomes. Neural network transformer models, initially developed for natural language processing (NLP) [72], offer a way to harness the temporal and longitudinal information in electronic health records (EHRs), often underutilised by traditional modelling approaches. The move toward generative pretrained transformers (GPTs) using autoregressive next-token prediction has resulted in large language models (LLMs) capable of zero- and few-shot performance across diverse tasks without task-specific retraining [12], inspiring analogous efforts in healthcare [34, 58, 32, 59, 63, 73]. Generative EHR models hold the promise of learning directly from rich, longitudinal data without relying on predefined rules, adapting to shifting epidemiology and capturing complex temporal associations that prove challenging for traditional models - properties which make them uniquely suited to modelling the complexity of COVID-19 pandemic. In applied use they offer the potential to support early detection, risk stratification, and simulation of clinical scenarios in a single model, without the need to retrain for each outcome of interest, enabling faster insights in a rapidly evolving public health emergency. In England, the Control of Patient Information (COPI) Regulations [42] provided a legal framework to make de-identified, routinely collected, national-scale NHS data available for COVID-19-related research [75]. These data include primary care, secondary care, COVID-19 testing, vaccination, and mortality data from the population of England (approximately 61 million people [56]) and are securely stored within the NHS England Secure Data Environment (NHSE SDE) [51]. Despite the potential of such a resource for the development and evaluation of AI models for COVID-19 research, computational restrictions within the NHSE SDE have limited previous projects to substantially smaller cohorts [20, 2]. In this work, we developed Foresight-England (Foresight-E), a 243-million-parameter transformer trained from scratch entirely within the NHSE SDE on linked primary and secondary care data from 54.9 million people (90%; total dataset size of 61 million). Foresight-E was designed for the zero-shot prediction of both the direct effects of COVID-19 (e.g., hospitalisation, mortality) and its indirect systemic impacts across the population. Access to this scale and demographic diversity is important to accurately model the COVID-19 pandemic, particularly to capture the outcomes of ethnic minority groups and patients with rare diseases or COVID-related complications who are only represented in statistically significant numbers at a population level [56, 70, 29]. Recent methodological research also shows that EHR foundation models exhibit scaling laws similar to those seen with LLMs, suggesting that larger scale training data improves performance [80]. Linking primary care records with secondary care, COVID-19 testing, vaccination data, and death registries provides a mechanism to capture the entire spectrum of the diseaseâfrom mild, community-managed presentations to critical hospital care. Crucially, this national-scale baseline allows for rigorous evaluation of algorithmic fairness across different ethnicities, age groups, and socioeconomic backgrounds. Indeed, most prior generative EHR models have been restricted to single healthcare institutions [59], specific EHR providers [73], or limited population subsets [34]. This limits their applicability to the English general population and risks algorithmic bias, a failure to maintain consistent performance across diverse demographic and clinical groups [1]. Furthermore, the true burden of COVID-19 extends far beyond acute viral infection. The pandemic caused profound, systemic disruptions to routine healthcare delivery, resulting in delayed diagnoses, altered treatment pathways, and the emergence of post-acute sequelae (such as Long COVID) which continue to impact both individuals and healthcare systems today. Whilst transformer models offer the potential to capture these changes during their training, the extent to which this generalises to future, unseen data remains unquantified. Therefore to rigorously evaluate a modelâs capacity to capture these critical indirect pandemic effects, and its ability to generalise to future data, it is necessary to test its predictive performance across the wider spectrum of human disease and beyond its training data, here encompassing over 1,400 clinical phenotypes in the year of 2023. As outlined in Project Status 2, quantitative results are unavailable, therefore we report the data pipeline, tokenisation, architecture, training, inference, and evaluation framework underpinning the model, assessed against the TRIPOD+AI [16] and PROBAST-AI [39] reporting guidelines (see Appendix C and D). We contribute the following: 1. Development of Foresight-E, the first national-scale generative foundation model of EHRs for COVID-19 research. 2. A reproducible methodology for longitudinal EHR tokenisation and model development. 3. A comprehensive evaluation framework to assess zero-shot prediction of COVID-19âs direct and indirect effects on a held-out disjoint test set. Comprising 6.1 million unique patients whose data was not used during model training, and temporally held out data. 4. First demonstration of a multi graphics processing unit (GPU)-accelerated foundation model pretrained on national-scale NHS data within the NHSE SDE. Together, these provide both a methodological blueprint for future EHR foundation models and a case study in the challenges of developing national-scale generative AI within secure health data environments. 4 Methods We developed Foresight-E for COVID-19 research using linked, de-identified, routinely collected national datasets [69, 75]. A key principle of Foresight-E is that the model is treated with the same security as the underlying data. Therefore, all processing, model training, inference, and evaluation occurred entirely within the âFive Safesâ framework of the NHSE SDE [51]. Neither the model weights nor any generated patient timelines can be exported outside this environment; only aggregated, non-disclosive evaluation metrics are eligible for release, via the approved Safe Output Service, subject to statistical disclosure control and review. Here we outline the datasets, patient timeline construction, tokenisation, model architecture, training regime, and our inference and evaluation strategy, including uncertainty estimation and comparative baselines. Each step was designed to support robust, generalisable, zero-shot medical event prediction for the direct and indirect effects of COVID-19. 4.1 Data To train and evaluate Foresight-E, longitudinal patient data were accessed from eight pseudonymised, linked, routinely collected national datasets covering 1 November 2018 to 31 December 2023 [69, 75]. These encompassed primary care (General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR) [48]), secondary care (Hospital Episode Statistics (HES): Outpatients (OP), Accident & Emergency (A&E), Admitted Patient Care (APC) and Critical Care (C) [49]), mortality (Office for National Statistics (ONS) Civil Registration of Deaths [21]), COVID-19 testing (UK Health Security Agency (UKHSA), formerly Public Health England (PHE), COVID-19 Second Generation Surveillance System (SGSS) [43]) and COVID-19 vaccination (NHS England COVID-19 Vaccination Status [44]) data (see Fig. 1(a)). The base cohort comprised individuals alive on or after 1 November 2019 (the GDPPR dataset inclusion criterion) with known age and sex, resident in England, GP-registered, and without conflicting death dates. To ensure a minimum one year of past medical history before predicting outcomes, we included data from 1 November 2018 onwards. Because primary care registration dates were unavailable to confirm exact prior follow-up lengths, this fixed calendar gap served as a population-level proxy for baseline history. To accommodate a dynamic cohort, individuals born after the study start date could enter the cohort at birth, however to enforce the minimum one year history for these newborns, we required them to reach at least one year of age prior to death or the end of the training period (31 December 2022). Ultimately, this yielded a nationally representative cohort of 61 million patients [56]. Table 1: PLACEHOLDER Selected characteristics of the training, validation, and test sets (used in Sections 5.1 and 5.2). Ages are computed at the time of the last input event in each patientâs timeline. Counts are rounded to the nearest five to comply with the NHSE SDEâs disclosure control requirements. Total sample sizes for each split are derived from previously reported counts [56]. The rationale for placeholders is provided in Section 5. Train Validation Test: 30-day COVID-19 Test: 1 Year Sex Female _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) Male _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) Ethnicity Asian _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) Black _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) Mixed _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) Other _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) Unknown _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) White _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) Age 0-9 _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) 10-19 _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) 20-29 _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) 30-39 _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) 40-49 _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) 50-59 _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) 60-69 _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) 70-79 _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) 80-89 _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) 90+ _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) _\_ (_\_%) Total 48.8M 6.1M 6.1M _\_ We created training (48.8M patients), validation (6.1M patients), and test sets (6.1M patients) via disjoint 10% patient samples. All 2023 events were reserved to test temporal generalisation, mimicking a prospective deployment on future, unseen data. Fig. 1 shows temporal coverage and partitioning of the datasets. (a) Linked primary/secondary care, deaths, testing, and vaccination datasets (1 Nov 2018â31 Dec 2023). (b) Disjoint 10% validation/test cohorts; 2023 held out for temporal generalisation evaluation. Figure 1: Datasets and cohort splits used in the training and evaluation of Foresight-E. Total sample sizes for each split are derived from previously reported counts [56]. 4.2 Patient Timelines Each patientâs history was encoded as a chronological sequence of dated, coded events, including diagnoses, procedures, medications, and other healthcare interactions. Clinical codes followed standard terminologies: ICD-10 [76] for HES diagnoses and ONS causes of death, OPCS-4 [47] for HES procedures, and SNOMED CT [50] for GDPPR records and COVID-19 vaccinations. We supplemented these with custom tokens for events not coded in standard terminologies, such as positive SARS-CoV-2 tests, hospital admission indicators, and primary or secondary diagnosis flags. No free text terms were available in the underlying data, or included in the model training. Because event timestamps were available only at day-level precision, we applied a consistent within-day ordering: COVID-19 vaccination; SARS-CoV-2 test; GDPPR events; HES OP; HES A&E; HES APC; HES C; and, lastly, death. Where relevant, within each dataset events were further ordered by admission date, then by primary diagnoses, secondary diagnoses, procedures, and finally alphanumerically. Geographic (region) and deprivation (Indices of Multiple Deprivation) data were excluded from the training data to avoid explicitly encoding these existing biases and enable held-out subgroup analysis [1]. Codes deemed sensitive (e.g. relating to sexually transmitted infections) were removed, using a clinically-informed codelist supplied by NHS England [41]. 4.3 Tokenisation We converted each patientâs event timeline into a sequence of discrete tokens representing clinical codes and temporal intervals. Each sequence began with static demographic tokens for sex (e.g., SEX_FEMALE) and ethnicity (e.g., ETHNICITY_ASIAN). We then added the sequence of clinical codes, ordered as described in Section 4.2, from each patientâs timeline as tokens. If two consecutive events occurred on different days, we inserted a time-difference token (e.g., TIME_DIFFERENCE_1) to represent the interval in days; no time-difference token was added between consecutive same-day events. To encode absolute calendar time, we inserted a YEAR_START token at the beginning of each calendar year, followed by an AGE_N token for the patientâs integer age (or AGE_UNBORN if not yet born). These yearly markers ensured that time gaps never exceeded 366 days and allowed the representation of age-related and absolute temporal patterns to generalise temporally. A lookup-based tokeniser (acting as a simple 1:1 dictionary) mapped all unique tokens in the training data to integer IDs, including clinical codes, age (integer years and unborn), 1â366 day gaps, and special tokens for padding (<PAD>) and sequence end (<EOS>). The resulting fixed tokeniser vocabulary contained approximately 40,000 tokens. At inference time, if a code was encountered that was not in the trained tokeniserâs vocabulary, it was dropped from the sequence without substitution, rather than using an unknown token (e.g., <UNK>). During training, we truncated each individual patientâs timeline to a maximum of 1,024 tokens using left truncation (discarding the earliest events) to retain the most recent clinical context. Truncation was applied at the event level followed by re-tokenisation, rather than truncating token sequences directly; this preserved demographic, year-start and age tokens and correctly recomputed time-difference tokens. For batching, we right-padded sequences to a uniform length with padding tokens, which were masked from attention and loss calculations. At inference, we pre-truncated inputs to 1,024âLforecast1,024-L_forecast, where LforecastL_forecast denotes the tokens allocated for forecast generation, instead of implementing dynamic on-GPU truncation. The maximum sequence length of 1,024 tokens was selected to balance capturing the maximum possible clinical context while maintaining adequate training and inference batch sizes on the available NVIDIA A10 GPUs. 4.4 Training Foresight-E is a 243-million-parameter transformer-decoder model trained to predict the next code in a tokenised EHR sequence, see Section A.1. This is analogous to how large language models (LLMs) are trained and can be considered as predicting what will happen next to a patient, based on their past medical history [12, 32, 59]. We adapted the open source Llama 2 architecture [71, 24] and trained the model from scratch with randomly initialised weights due to the inability to import pretrained models into the NHS SDE and the use of a custom vocabulary, with clinical codes and custom tokens, rather than natural language. Training comprised a single pass through the dataset with periodic checkpointing; the checkpoint with the lowest validation loss was selected as the final model. Hyperparameters followed established defaults [32, 3]. Complete architectural and training details are provided in Section A. Figure 2: Schematic of next token prediction training of Foresight-E on a synthetic patient timeline. Tokens at position IDs 2â7 are omitted for clarity. Given the sequence of tokens up to position i, the model predicts the probability distribution over the vocabulary for the token at position i+1i+1. These probabilities are compared with the true next token to compute the cross-entropy loss for self-supervised training. Note that the initial static demographic tokens are masked ([Ignore]) during the loss calculation to prevent the model from learning to predict sequence progression solely from demographic features. The final and padding tokens are also ignored because they have no label to compare against. 4.5 Zero-Shot Inference As a generative model, once trained, Foresight-E can predict future clinical events across the breadth of its vocabulary from patient histories, without requiring task-specific supervised fine-tuning (âzero-shotâ inference). This can be viewed as analogous to LLMs âautocompletingâ a sentence. At inference, the model is prompted with a tokenised patient timeline (see Section 4.3) and produces a probability distribution over the next token. We sample one token from this distribution and append it to the input. This autoregressive procedure is repeated until a prespecified stopping criterion is met: a sentinel event token (e.g., mortality), a maximum forecast horizon (e.g., 30 days or 1 year), or a limit on newly generated tokens (e.g., 300). The generated tokens are then decoded to form the predicted event timeline, see Figure 3 Figure 3: Schematic of idealised zero-shot inference on a synthetic patient timeline. Given a history ending with a positive SARS-CoV-2 test, Foresight-E predicts a hospital admission within the next 30 days. Tokens 2â7 are omitted for clarity. Because greedy decoding yielded repetitive sequences, we used multinomial sampling from the modelâs output distribution to generate diverse and clinically plausible trajectories. We did not use parameter-dependent sampling techniques (e.g., top-k, temperature), which would have introduced additional task-specific hyperparameters. A key-value (KV) cache was used to increase generation throughput by avoiding the recomputation of past steps. To quantify predictive uncertainty at the individual patient level, an essential consideration for safe and reliable clinical decision-making [60], we performed S=48S=48 independent stochastic rollouts per patient. This was the maximum batch size possible on a single A10 GPU [52] given Foresight-Eâs size and context length. Inference was parallelised across eight GPUs for throughput [6]. For event i and horizon t (in days), the predicted probability was estimated as [59] PâĄ(Eventi,t)=si,tS,P(Event_i,t)= s_i,tS, (1) where si,ts_i,t denotes the number of rollouts in which the event occurred within the horizon. This formulation provides probability estimates across all medical events observed during training, enabling probabilistic zero-shot predictions across diverse clinical outcomes. Figure 4: Visualisation of an idealised Foresight-E zero-shot inference for the task of 30-day hospitalisation following a positive SARS-CoV-2 test. A tokenised input timeline, shortened for visualisation, is used to prompt three independent predicted timelines in this example. âTime Diffâ tokens are in units of days. In two of the predicted timelines, hospitalisation is predicted within the next 30 days (at 20 and 25 days, respectively). The remaining timeline predicts the next event at 100 days, falling outside the target window. Therefore, the predicted probability is 2/3. 4.6 Evaluation We evaluated Foresight-E on two acute COVID outcomes: hospitalisation and death within 30 days following a positive SARS-CoV-2 test. Eligible patients in the test set had at least one positive test between 23 January 2020, the date of the first recorded COVID-19 case in the UK, and 1 December 2023 [35]. While mass community testing in England ended on 1 April 2022, we deliberately included tests beyond this date [13]. Although testing post-April 2022 is inherently selective, capturing predominantly high-risk, hospitalised, or healthcare worker populations, evaluating our model across this policy shift assesses its robustness to changing clinical testing practices and population risk profiles across the pandemic. Timelines were truncated after the first positive test, and predictions generated for the next 30 days. Hospitalisation was defined as an admission in HES APC (APC token), and death as an entry in the Office for National Statistics (ONS) death registry (DEATH token). Inference was completed in 2 days of wall-clock compute time. Secondly, we assessed the âindirectâ effects of COVID-19 in the held-out year of 2023, evaluating emergency hospitalisation, all-cause mortality, and the onset of over 1,400 Phecode-defined diseases. We evaluated these outcomes across the entire held-out cohort, irrespective of a formally recorded positive SARS-CoV-2 test or COVID-19 diagnosis. This population-wide evaluation accounts for the high levels of prior SARS-CoV-2 infection by 2023, and captures the systemic healthcare disruptions that affected patients across all care pathways [8, 7, 40, 53]. Here, timelines were truncated after the 2023 year-start token and subsequent events predicted over a one-year horizon. Emergency admissions were identified via admission method codes in HES APC, and mortality as above. Phenome-wide disease onset was defined by the first occurrence of each Phecode, a previously validated collection of over 1,400 groups of ICD-10 codes (example shown in Table B) [78], recorded in HES APC. Inference was completed in 15 days of wall-clock compute time. 4.6.1 Metrics For evaluation, each outcome was treated as a binary classification task: given a patientâs history, the model provided an estimate of the probability of event i occurring within t days, and this was compared against the observed outcome. This allowed for the evaluation of single medical event prediction as well as bespoke phenotypes composed of multiple clinical codes or more complex definitions. Under this binary classification formulation, false positives and false negatives are quantified and correspond to the concepts of hallucinations and omissions in LLMs, respectively. Competing risks (such as death prior to the onset of an outcome) were not explicitly modelled as separate states; instead, the prediction of a death token terminated the sequence generation. Discriminative performance was assessed using the area under the receiver operating characteristic curve (AUROC) [28] and the area under the precision-recall curve (AUPRC) [82], the latter being commonly used under class imbalance [54, 61]. Overall performance and calibration were jointly measured using the Brier score [9]. We also generated ROC, precision-recall, and calibration curves for visual inspection. To reflect potential use in a clinical screening workflow, we also evaluated recall at a fixed 10% false-positive rate (FPR10), consistent with prior studies [15] and quantifying performance at a hypothetical operational threshold. Confidence intervals (95%) were obtained via non-parametric patient-level bootstrapping with the percentile method, using 1,000 resamples for all analyses except the Phecode tasks (where 100 were used due to computational cost). 4.6.2 Subgroup Analyses We conducted subgroup analyses of mortality predictions at both the 30-day postâSARS-CoV-2âpositive test and one-year horizons, stratifying test-set patients independently by age, sex, ethnicity, and vaccination status. All of these stratification variables were tokenised and included in the modelâs training data; age and vaccination status were embedded longitudinally within the patient timelines, while sex and ethnicity were included as static demographic tokens at the start of each timeline. For each subgroup, we calculated the AUROC and AUPRC in comparison with the overall cohort. Additionally, we investigated how performance on these metrics varied with the number of historical patient events provided as model input. Finally, we assessed changes in AUROC and AUPRC relative to the timing of the positive SARS-CoV-2 test to determine whether the Foresight-E model could model the changing epidemiology of COVID across the pandemic. 4.7 Baseline Methods We benchmarked Foresight-E against supervised classifiers trained separately for all-cause mortality and hospitalisation at 30 days after a positive SARS-CoV-2 test, and for the same outcomes over a one-year horizon in 2023. These comparators used the same training split as Foresight-E but were task-specific, in contrast to Foresight-Eâs zero-shot capability. The first was a simple logistic regression model [62] using age, sex, and ethnicity (one-hot encoded, meaning each categorical variable was represented as a binary vector where only the position corresponding to the observed category is set to one), with vaccination status added for acute SARS-CoV-2 outcomes. The second was an XGBoost model [19] using count vectors of all medical codes from the Foresight-E vocabulary, providing equivalent structured clinical data but without temporal information. To allow training outcomes to be observed, input cut-offs were set to 1 December 2022 for the 30-day tasks and 1 January 2022 for the one-year tasks. This requirement highlights a key advantage of Foresight-Eâs self-supervised learning: it is trained without requiring explicitly defined outcome labels. This eliminates the need to withhold a temporally separated subset of data for outcome labelling. Performance for all models was evaluated using AUROC, AUPRC, and Brier score, as described in Section 4.6.1. 4.8 Governance The North East - Newcastle and North Tyneside 2 research ethics committee provided ethical approval for the CVD-COVID-UK/COVID-IMPACT research program (REC No 20/NE/0161) for approved research projects to access, within secure trusted research environments, whole-population, de-identified data from EHRs collected as part of patientsâ routine healthcare. The CVD-COVID-UK/COVID-IMPACT programme, led by the BHF Data Science Centre [11], received approval to access data in the NHSE SDE service for England from the Independent Group Advising on the Release of Data (IGARD) [46] via an application made in the Data Access Request Service (DARS) Online system (ref. DARS-NIC-381078-Y9C5K) [45]. The CVD-COVID-UK/COVID-IMPACT Approvals & Oversight Board that includes patient and public advisors [10] subsequently granted approval to this project (CCU078: Foresight: a generative AI model of patient trajectories across the COVID-19 pandemic) in December 2023 to access the data within the NHSE SDE service for England. Patient and public involvement was included in the approvals process and has continued to shape the research and communications through Patient and Public Involvement and Engagement sessions organised via British Heart Foundation (BHF) Data Science Centre [10]. 4.8.1 Data Availability The data used in this study are available in the NHSE SDE service for England, but as restrictions apply, they are not publicly available [51]. The de-identified data used in this study were made available to accredited and approved researchers only. Those wishing to gain access to the data should contact bhfdsc@hdruk.ac.uk in the first instance, noting that as detailed in Project Status 2 data access is currently paused. 4.8.2 Code Availability All data preparation, model training, and evaluation code was intended to be released, following export and review via the SDEâs Safe Output Service, on GitHub at: https://github.com/BHFDSC/CCU078_Foresight-England, including a requirements.txt specifying all package versions. However as detailed in Project Status 2 data access is currently paused meaning code cannot be exported and shared. The code used to produce the placeholder results figures was developed outside the SDE, without access to the underlying data, and is available at https://github.com/simonEllershaw/foresight_placeholder_graphs Due to data restrictions, trained model weights and artefacts are only accessible to a subset of approved CVD-COVID-UK/COVID-IMPACT consortium researchers on a dedicated Foresight-E cluster within the NHSE SDE. Those wishing to gain access to the data should contact bhfdsc@hdruk.ac.uk in the first instance, noting that as detailed in Project Status 2 data access (including model weights, artefacts and code) is currently paused. All analyses were executed within the NHSE SDE [51] using Databricks Runtime 14.3 LTS for ML [18]; training/evaluation used an AWS g5.48xlarge instance with eight NVIDIA A10 GPUs [6, 52]. AWS and Databricks had no access to the underlying datasets or trained AI model and no control over the research or its findings. 5 Results As detailed in Project Status 2, an initial quantitative evaluation has been completed, but the ongoing pause in data access prevents export of the results from the SDE. To be transparent about what was evaluated and what outputs were intended for publication, we present the relevant tables and figures with placeholder values, marked with â_\_â, and placeholder figures where data cannot be shown. 5.1 30-day COVID-19 Mortality and Hospitalisation Prediction To demonstrate the utility of Foresight-E in predicting the direct effects of COVID-19 without fine-tuning, we applied it to the important clinical prediction task of forecasting 30-day mortality and hospitalisation following a positive SARS-CoV-2 test. Foresight-E achieved an AUROC of _\_ (_\_-_\_, 95% CI) and _\_ (_\_-_\_, 95% CI) for 30-day prediction of mortality and hospitalisation, respectively. For comparison, the baseline supervised logistic regression and XGBoost models yielded AUROCs of _\_ and _\_, and _\_ and _\_, respectively. Calibration and discrimination was quantified using Brier scores, yielding _\_ and _\_ for the mortality and hospitalisation endpoints, respectively. At a 10% false positive rate (FPR), detection rate (DR10) was _\_ and _\_, respectively. ROC, precision-recall, and calibration curves for each task are shown in Fig. 5, with 95% confidence intervals estimated via 1,000 sampled bootstrap iterations. (a) Mortality (b) Hospitalisation Figure 5: PLACEHOLDER. ROC, PR, and calibration curves showing Foresight-Eâs prediction performance for 30-day mortality and hospitalisation following a positive SARS-CoV-2 test. Shaded regions represent 95% confidence intervals estimated via 1,000 bootstrap iterations. Counts are rounded to the nearest five to comply with the NHSE SDEâs disclosure control requirements. 5.1.1 COVID-19 Outcomes Across the Pandemic To evaluate Foresightâs ability to model the changing epidemiology of COVID-19 across the duration of the pandemic (including varying case rates, mortality, and circulating viral variants), we performed a subgroup analysis stratified by the calendar year and month of the patientâs first positive SARS-CoV-2 test. Across these temporal strata, we observed AUROC values ranging between _\_ and _\_. On the 2023 data, a temporal holdout set excluded from training or validation, the computed AUROC was _\_. Figure 6: PLACEHOLDER. Temporal evaluation of 30-day mortality prediction across the COVID-19 pandemic. Monthly performance of Foresight-E in predicting death within 30 days of a positive SARS-CoV-2 test, measured by AUROC (top panel) and AUPRC (second panel). Error bars represent 95%95\% confidence intervals estimated via 1,000 bootstrap iterations. Also shown are monthly mortality occurrence rates (third panel) and the number of positive SARS-CoV-2 tests (bottom panel), illustrating changes in disease dynamics over time. Shaded background regions denote the predominant SARS-CoV-2 variants [74] with labels shown along the upper x-axis. Counts are rounded to the nearest five to comply with the NHSE SDEâs disclosure control requirements. 5.1.2 COVID-19 Outcomes by Subgroup Fig. 7 shows the performance of the Foresight-E model in predicting 30-day mortality following a positive SARS-CoV-2 test, stratified by age, ethnicity, sex, and COVID-19 vaccination status. To assess for performance disparities across demographic groups, we computed AUROC scores of _\_ for White patients, _\_ for Asian patients, _\_ for Black patients, and _\_ for Mixed/Other ethnicities. Between sexes, AUROC scores were _\_ for male patients and _\_ for female patients. Across age brackets, AUROC values measured _\_ to _\_, mapped against a crude mortality rate of _\_% to _\_% in those respective demographics. Figure 7: PLACEHOLDER. Variation in AUROC and AUPRC metrics for the Foresight-E model predicting 30-day mortality after a positive SARS-CoV-2 test, stratified by age, ethnicity, sex, and COVID-19 vaccination status. Error bars show 95% confidence intervals calculated by 1,000 bootstrap iterations. Also shown is the mortality occurrence rate, as well as the count of patients for each group. Performance is consistent across ethnicity and sex, but varies notably by age group and vaccination status. Counts are rounded to the nearest five to comply with the NHSE SDEâs disclosure control requirements. 5.1.3 COVID-19 Outcomes By Event History Fig. 8 plots model performance predicting 30-day mortality after a positive SARS-CoV-2 test against the number of historical patient events provided as input to Foresight-E. We measured a correlation of _\_ between the number of past medical events and AUROC, contextualised by a correlation of _\_ between the number of past medical events and the outcome. The distribution of the number of past events across the cohort exhibited a skewness of _\_. Figure 8: PLACEHOLDER. Variation in AUROC and AUPRC metrics for the Foresight-E model predicting 30-day mortality after a positive SARS-CoV-2 test with the number of past medical events recorded in a patientâs timeline. Also shown are the binned frequency of past events and corresponding mortality occurrence rates. A bin size of 16 events was used, and the final bin includes all patients with more than 498 events. Counts are rounded to the nearest five to comply with the NHSE SDEâs disclosure control requirements. 5.2 Predicting Indirect Effects of COVID-19 in 2023 To evaluate Foresight-Eâs ability to predict the indirect effects of the COVID-19 pandemic, we simulated a prospective deployment. All patient timelines were truncated at 1 January 2023, using the YEAR_START token. Foresight-E was not exposed to any event data from 2023 during training, thereby testing generalisation to unseen future events. 5.2.1 Emergency Hospitalisation and Mortality Prediction The ROC, precision-recall, and calibration curves are presented in Fig. 9. Foresight-E achieved an AUROC of _\_ (_\_-_\_, 95% CI) and _\_ (_\_â_\_, 95% CI) for emergency hospitalisation and mortality, respectively. For comparison, the baseline supervised logistic regression and XGBoost models yielded AUROCs of _\_ and _\_, and _\_ and _\_, respectively. Calibration and discrimination was quantified using Brier scores, yielding _\_ and _\_ for emergency hospitalisation and mortality endpoints, respectively. (a) One-year mortality prediction: ROC, PR, and Calibration Curves. (b) One-year emergency hospitalisation prediction: ROC, PR, and Calibration Curves. Figure 9: PLACEHOLDER. Prediction performance of Foresight-E in 2023 using ROC, PR, and calibration curves. Shaded regions represent 95% confidence intervals estimated via 1,000 bootstrap iterations. Counts are rounded to the nearest five to comply with the NHSE SDEâs disclosure control requirements. 5.2.2 Phecode Predictions To assess the breadth of Foresight-Eâs capacity to predict the indirect effects of COVID, we evaluated performance on over 1,400 distinct Phecodes, previously validated groupings of ICD codes representing clinically meaningful phenotypes [78]. AUROC values varied _\_-_\_ across Phecodes, with an _\_ association between event prevalence and predictive performance (Fig. 10). Figure 10: PLACEHOLDER. Relationship between AUROC and event occurrence rate for Foresight-Eâs 1-year prediction in 2023. _\_ AUROCs were observed for more _\_ Phecodes, such as _\_. 6 Discussion We developed Foresight-E, a 243-million-parameter generative foundation model trained de novo on longitudinal electronic health records from 54.9 million patients, entirely within the NHS England Secure Data Environment. Designed as a research pilot for COVID-19, the model was evaluated on its zero-shot capability to predict both the direct acute outcomes of SARS-CoV-2 infection and the pandemicâs enduring, system-wide indirect effects. By successfully engineering this pipeline within existing national infrastructure, this work establishes a methodological template for future pandemic response and broader applications, from clinical risk prediction to population-level healthcare planning. As quantitative results on our 6.1-million-patient test set are currently withheld pending an ongoing governance review, we focus our discussion below on the core methodological strengths and limitations of this work, including data integration, modelling and evaluation, before discussing future directions for the field. 6.1 Data Foresight-Eâs key strength is its first use of national-scale, routinely collected EHRs spanning primary care, secondary care, COVID-19, and death registrations for the development of a generative AI model. This ensures Foresight-E is trained on diverse data representative of the general population [56], helping to mitigate algorithmic bias [5], enabling prediction of rare events [70, 68], and generating evidence for the methodologyâs translational potential through aligning training data with intended use populations. Given that primary care accounts for most healthcare delivery, integrating this data provides a more complete representation of an individualâs health, enables earlier risk stratification, and may ultimately guide preventive interventions to avert disease onset, complications and costly secondary care. Despite its breadth, the datasets reflect typical challenges of routinely collected health data: incomplete or imprecise coding [31], historical biases in care [64, 5], and shifts in recording practices [79], especially during the COVID-19 pandemic [55]. Notably, the GDPPR primary care dataset contains a subset of codes, specifically, those already available from previous GPES extracts that were deemed relevant for COVID-related research. As a result, it omits examples of both common and rare diseases, as well as signs and symptoms data that is crucial for understanding disease presentation and evolution - for example, COVID presenting with anosmia, or non-specific symptoms of long COVID, such as fatigue. Furthermore, GDPPR only contains those alive on or after 1 November 2019. This means that individuals who died before this date are excluded, restricting the ability to model or evaluate long-term disease progression or events in the pre-COVID era. While this is not a limitation for models focused on COVID-19 infection, since such cases did not exist before early 2020, it does reduce the available historical context. Expanding the temporal window to include earlier records would be expected to increase predictive performance and improve the capacity to model long-term outcomes, but would require methods capable of efficiently handling longer sequences [77] and risk issues such as immortal time-bias without changes in dataset inclusion criteria. While Foresight-E is strictly confined to the structured data approved for COVID-19 research, future applications of this methodological approach could be enhanced by incorporating additional modalities such as clinical notes or medical imaging. The development of foundational multi-modal transformer models in the general domain shows how the methods presented here could be extended [57, 25]. However, such data is not currently available at a national scale, and provisioning new data streams for future models would require appropriate governance, infrastructure, and technical methods, including de-identification and pre-processing. 6.2 Model Foresight-E uses a 243-million-parameter Llama 2-style transformer decoder. While larger than most prior EHR models [63] and trained on national scale data, the scale of data, model size, and compute is still orders of magnitude smaller than the general domain [38]. Foresight-E is designed to be data-driven, rather than relying on predefined parametric models such as exponential hazardâstyle structures [63, 65]. Although we trained from scratch due to a custom vocabulary and NHSE SDE import restrictions, smaller-scale studies indicate that starting with an LLM pre-trained on vast general data and then fine-tuning for medical-event prediction is beneficial to performance [33, 22]. If future NHSE SDE policy permits importing pretrained models, initialising from a general model, and adapting in-domain is a promising direction. Our tokenisation strategy preserves the recorded granularity of clinical codes and day-level timing of the EHR data , rather than aggregating codes into higher level categories [63] or phenotypes [34], filtering low-occurrence events [32] or binning time [59]. This maximised diagnostic specificity and allowed prediction of rare events often excluded in other models [68]. However, this enlarged the vocabulary and prediction space and reduced per-token training frequency. Furthermore, all event codes present in the training set were added to the tokeniserâs vocabulary. Therefore, during inference, codes absent from the training data could not be represented because the vocabulary did not include an unknown token (e.g., <UNK>), so we excluded them, discarding potentially clinically relevant information. The tokeniser could be further improved by exploiting the hierarchical structure of clinical ontologies [59], for which there is evidence of performance gains on EHR prediction tasks [67]. Building on previous work that models chronological age and relative time between events, we additionally encoded absolute calendar time to model the changing healthcare system during the COVID pandemic. Because Foresight-E has a maximum context length of 1,024 tokens, long histories required truncation. NaĂŻve token-level truncation (e.g., dropping the first n tokens) corrupts timelines by removing demographic, year-start, or age tokens or invalidating time-difference tokens. We therefore applied event-level truncation followed by re-tokenisation as a pre-processing step for training. During autoregressive inference, however, all operations must run on the GPU, making event-level truncation infeasible. We therefore used fixed maximum input and generation lengths, which occasionally discarded more tokens than necessary and, in some cases, prevented forecasted trajectories from reaching the final time horizon (e.g. 1 year). Previous work [59] constructs training examples by concatenating all patient EHRs into a single sequence and then randomly sampling subsequences of a fixed length. Although this maximises computational efficiency by eliminating the need for padding tokens, it allows individual training examples to span across multiple patients. This risks the model learning spurious cross-patient relationships. To avoid this, we adopt a patient boundary-preserving strategy [80], forming sequences strictly at the patient level, which comes at the expense of computational efficiency. Future work could incorporate intra-document causal masking [81] to enable efficient concatenation without information leakage. 6.3 Inference We extended prior probabilistic patient-trajectory approaches [59], applying Foresight-E in a zero-shot setting to forecast over 1,400 medical events, across 30-day and 1-year time horizons relevant to direct and indirect COVID-19 outcomes. This offers the broad ability to predict differential diagnoses and clinical trajectories, rather than single outcomes, but is computationally intensive, particularly at a population scale. Comparative studies with task-specific fine-tuning are needed to clarify efficiencyâflexibility trade-offs. Our current setup generates 48 trajectories per patient, constrained practically by the NVIDIA A10 GPUâs memory. Exploring how probability estimates stabilise with the number of samples and alternative decoding strategies (beam search, top-k, temperature) may yield gains in efficiency, accuracy, and calibration, but would require tuning. Furthermore, quantifying the temporal horizon over which Foresight-E can reliably forecast patient outcomes, both in absolute time and in terms of tokens, requires further study. 6.4 Evaluation We evaluated Foresight-Eâs ability to predict COVID-19 outcomes in two settings designed to reflect real-world deployment challenges. First, we assessed its ability to predict direct COVID-19 outcomes during the pandemic, a period marked by shifting conditions such as emerging viral variants, changing testing protocols, public health interventions and population immunity. Unlike earlier models [59], Foresight-E explicitly encodes both absolute calendar time and patient age, enabling it to adapt to these temporal shifts. Forecasting future novel threats, outside the current training data and vocabulary, could be supported by continued pretraining with vocabulary expansion on the latest batches of newly collected data. Second, we sought to predict the indirect effects of COVID-19 and simulate a prospective deployment by predicting events in 2023, one year beyond the training period, on over 1,400 Phecodes, emergency hospitalisation, and all-cause mortality. COVID-19 highlights the challenges of temporal data shifts. Foresight-E was trained during the height of the COVID-19 pandemic, a period of profound healthcare system disruption and excess all-cause mortality, which left an enduring and evolving legacy in the evaluation period of 2023, encompassing the indirect effects the pandemic exerted on individuals, healthcare systems, and society at large [7]. This represents, to our knowledge, the broadest zero-shot evaluation of an EHR foundation model to date. Working at the scale of the English population meant facing a low occurrence rate for many acute outcomes, in contrast to models trained solely on high-acuity inpatient cohorts (e.g., MIMIC-IV [27]). While this sparsity poses challenges, it also demonstrates Foresight-Eâs potential utility for COVID-specific population screening as well as high-risk patient monitoring, such as identifying vulnerable individuals to prioritise targeted interventions (e.g. vaccination, antivirals), by training on both general âhealthyâ and acutely ill patient timelines. In order to critically evaluate Foresight-Eâs potential for population screening, we additionally calculated recall at a fixed 10% false-positive rate, consistent with prior studies [15] and representing a hypothetical operational threshold. Whilst evaluating each patient trajectory as multiple binary prediction tasks allowed use of established metrics (AUROC, AUPRC, Brier score) to quantify performance on clinically-relevant tasks, these fail to capture the overall fidelity of generated trajectories, an important direction for future work. A key limitation of this work is the inability to benchmark against commonly used clinical risk prediction tools, due to a lack of required data (e.g. 4C Mortality Score for COVID requires physiological measurements and test results [30]), limited data duration (e.g. QRisk measures 10-year cardiovascular disease risk [23]) and the fact that where risk scores were recorded the resulting outcomes were then conditioned on resultant clinical decision making [37]. For methodological comparators NHSE SDE constraints meant external pretrained models could not be imported and resource limitations restricted the number of comparator models trained. External validation, to assess model generalisability across different populations, was not possible as Foresight-E was developed entirely inside the NHSE SDE and the trained model could not be transferred out of the environment (including to another SDE), nor could other cohorts be imported into the NHSE SDE. Explainability techniques, such as attention-weighted visualisation, gradient-based saliency, or counterfactual generation, are left as directions for future work but could help clinicians understand the modelâs forecasts by scrutinising learned associations for known or plausible patterns, as well as the presence of spurious correlations, such as shortcut learning. Such methods could support safe adoption in practice by explaining why the model anticipates particular outcomes, as well as identifying potentially modifiable risk factors as targets for interventions to optimise health. 6.5 Future Directions and Considerations Foresight-E was developed as a research pilot strictly for COVID-19-related research and is not a validated clinical tool. The model is therefore confined to this scope, and any future directions for the methodology are entirely contingent on navigating the significant challenges outlined below. The generative, zero-shot forecasting methodology demonstrated by models like Foresight-E offers broad potential, including forecasting population health, stratifying groups at increased risk of adverse outcomes, and enabling personalised risk prediction to guide preventive interventions. Beyond direct clinical care, potential applications include improving clinical trial efficiency through prognostic enrichment or advancing drug discovery by better modelling disease trajectories. Models like Foresight are by design associative, rather than causal, and therefore a priority area for research is the robust evaluation of the extent to which counterfactual questions can be answered, a crucial step toward creating robust digital twins and enabling trustworthy in-silico trials. Moving from research to application would require secure, real-time model deployment, a capability beyond the current NHSE SDE infrastructure. This gap is particularly critical for large generative models, which can inadvertently memorise and expose sensitive training data [14]. A potential mitigation strategy is to deploy such models within a secure environment behind narrowly scoped APIs. These would provide only predefined, validated outcomes, such as calibrated risk scores, rather than open-ended generative trajectories, thereby constraining vectors for data extraction and mitigating privacy risks [26]. However, these ambitions are secondary to the fundamental legal and governance challenges. The most significant barrier to any extension of this methodology is that there is currently no lawful basis to use this national dataset beyond emergency directions issued specifically for COVID-19 pandemic research [42]. Therefore, advancing this work requires new legal permissions through transparent public consultation, paired with a clear public benefit. A lawful basis and a social licence must go hand in hand; this requires deep and sustained engagement with patients, the public, and professional bodies. Furthermore, the path from a research model to a trustworthy clinical tool would necessitate additional rigorous evaluation and a clear route to regulatory approval as a medical device. Addressing these socio-technical challenges is the central prerequisite for future progress in this domain. 7 Conclusion We have presented Foresight-E, a 243-million-parameter transformer trained on national-scale EHR data from 54.9 million NHS patients and evaluated on its ability to perform zero-shot prediction across âź 1.4k COVID-related outcomes for a 6.1-million-patient test set. We outline the data integration pipeline, tokenisation strategy, model architecture, training procedure, and inference and evaluation framework. Although quantitative results are currently withheld pending ongoing discussions, this work demonstrates that it is technically feasible to develop a foundation model for healthcare entirely within existing NHS infrastructure. By combining routinely collected population-scale EHRs with modern generative modelling, Foresight-E offers a blueprint for zero-shot healthcare AI systems. Rebuilding beyond COVID-restricted datasets, expanding access to broader clinical modalities, and developing safe deployment pathways could enable models like Foresight-E to support both population-level planning and individualised care. Realising this potential will require not only technical advances but also transparent governance, sustained public and professional engagement, and rigorous evaluation in real-world clinical settings to generate evidence for regulatory approval. 8 Acknowledgements The British Heart Foundation Data Science Centre (grant No SP/19/3/34678, awarded to Health Data Research UK) funded co-development (with NHS England) of the SDE service for England, provision of linked datasets, data access, user software licences, computational usage, and data management and wrangling support, with additional contributions from the HDR UK Data and Connectivity component of the UK Government Chief Scientific Adviserâs National Core Studies program to coordinate national COVID-19 priority research. Consortium partner organisations funded the time of contributing data analysts, biostatisticians, epidemiologists, and clinicians. AWS provided the compute credits which made this work possible. Databricks provided technical support. AWS and Databricks had no access to the underlying datasets or trained AI model and no control over the research or its findings. SE and CT are funded by the UKRI Centre for Doctoral Training in AI-enabled healthcare (EP/S021612/1). CT also receives support from a Kingâs College London AI+ Senior Academic Fellowship, a MRC Clinical Top-Up, NIHR Biomedical Research Centre at UCL Hospital NHS Trust, and Health Data Research UK. AMW is supported by the BHF Data Science Centre (HDRUK2023.0239), Health Data Research UK (Big Data for Complex Disease-HDR-23012), and as an NIHR Research Professor (NIHR303137). Her research is also supported by core funding from the British Heart Foundation (RG/F/23/110103), NIHR Cambridge Biomedical Research Centre (NIHR203312) [*], BHF Chair Award (CH/12/2/29428), Cambridge BHF Centre of Research Excellence (RE/24/130011), and by Health Data Research UK (HDRUK2023.0028), which is funded by the UK Medical Research Council, Engineering and Physical Sciences Research Council, Economic and Social Research Council, Department of Health and Social Care (England), Chief Scientist Office of the Scottish Government Health and Social Care Directorates, Health and Social Care Research and Development Division (Welsh Government), Public Health Agency (Northern Ireland), British Heart Foundation and the Wellcome Trust. RDâs work is supported by (1) National Institute for Health Research (NIHR) Biomedical Research Centre at South London and Maudsley NHS Foundation Trust and Kingâs College London. (2) Health Data Research UK, which is funded by the UK Medical Research Council, Engineering and Physical Sciences Research Council, Economic and Social Research Council, Department of Health and Social Care (England), Chief Scientist Office of the Scottish Government Health and Social Care Directorates, Health and Social Care Research and Development Division (Welsh Government), Public Health Agency (Northern Ireland), British Heart Foundation and Wellcome Trust. (3) The National Institute for Health Research University College London Hospitals Biomedical Research Centre. This work was carried out with the support of the BHF Data Science Centre led by HDR UK (BHF Grant no. SP/19/3/34678). This study made use of de-identified data held in the NHSE SDE service for England and made available via the BHF Data Science Centreâs CVD-COVID-UK/COVID-IMPACT consortium. This work used data provided by patients and collected by the NHS as part of their care and support. We would also like to acknowledge all data providers who make health-relevant data available for research. We thank partners at NHS England, AWS, Databricks, and the BHF Data Science Centre for their support, without which this project would not have been possible. We extend our sincere gratitude to the public contributors of the British Heart Foundation Data Science Centre for providing critical review, stimulating discussion, and constructive feedback throughout the projectâs lifecycle, including their invaluable guidance in shaping the lay summary. 9 Contributor Statement S.E. and C.T. contributed equally to this work and share joint first authorship. SE: Conceptualisation, Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualisation, Writing â original draft, Writing â review & editing. CT: Conceptualisation, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualisation, Writing â original draft, Writing â review & editing. ZK: Methodology, Writing â review & editing. SD: Methodology, Writing â review & editing. H: Writing â review & editing. CS: Funding acquisition, Resources, Writing â review & editing. AMW: Writing â review & editing. AS: Writing â review & editing. RD: Conceptualisation, Funding acquisition, Methodology, Resources, Supervision, Writing â review & editing. In strict accordance with the information governance and data security protocols of the NHS England Secure Data Environment (SDE), access to the underlying datasets and dedicated Foresight cluster was restricted to only CT and SE. Consequently, all data curation, model training, formal data analysis, software implementation, and quantitative evaluation were conducted exclusively by CT and SE within the secure environment. CT is the guarantor for this work. CS was previously the Director of the British Heart Foundation (BHF) Data Science Centre and coordinated approvals for and access to data within the NHS Digital Trusted Research Environment for England for CVD-COVID-UK/COVID-IMPACT. 10 Disclosures SE contracted part-time for Parexel International during the period this work was conducted. CT was previously employed by LifeArc, and has received research funding via the UCL-GSK Phenomics Hub from GSK. ZK is a co-founder of Nuraxi. RD is a co-founder of CogStack and Onsentia. ADS receives research funding from BMJ Publishing Group. The remaining authors declare no competing interests. None of these commercial organisations had any involvement in the funding, study design, data access, model training, evaluation, or execution of the Foresight-England project, nor do they have access to the underlying data, code, or model weights. References [1] J. E. Alderman, J. Palmer, E. Laws, M. D. McCradden, J. Ordish, M. Ghassemi, S. R. Pfohl, N. Rostamzadeh, H. Cole-Lewis, B. Glocker, et al. (2025) Tackling algorithmic bias and promoting transparency in health datasets: the standing together consensus recommendations. The Lancet Digital Health 7 (1), p. e64âe88. Cited by: §3, §4.2. [2] F. Allery, M. Pineda-MoncusĂ, C. Tomlinson, N. Pontikos, J. H. Thygesen, S. Khalid, and CVD-COVID-UK/COVID-IMPACT Consortium (2023) Towards mitigating health inequity via machine learning: a nationwide cohort study to develop and validate ethnicity-specific models for prediction of cardiovascular disease risk in COVID-19 patients. medRxiv. Cited by: §3. [3] Andrej Karpathy (2025) nanoGPT. Note: https://github.com/karpathy/nanoGPTAccessed: 2025-04-15 Cited by: §A.3, §4.4. [4] S. Armstrong (2025) NHS england faces investigation over granting foresight access to gp patient data. British Medical Journal Publishing Group. Cited by: §2. [5] A. Arora, J. E. Alderman, J. Palmer, S. Ganapathi, E. Laws, M. D. Mccradden, L. Oakden-Rayner, S. R. Pfohl, M. Ghassemi, F. Mckay, et al. (2023) The value of standards for health datasets in artificial intelligence-based applications. Nature medicine 29 (11), p. 2929â2938. Cited by: §6.1, §6.1. [6] AWS (2025) Amazon EC2 G5 Instances. Note: https://aws.amazon.com/ec2/instance-types/g5/Accessed: 2025-04-14 Cited by: §4.5, §4.8.2. [7] S. Ball, A. Banerjee, C. Berry, J. R. Boyle, B. Bray, W. Bradlow, A. Chaudhry, R. Crawley, J. Danesh, A. Denniston, et al. (2020) Monitoring indirect impact of COVID-19 pandemic on services for cardiovascular diseases in the uk. Heart 106 (24), p. 1890â1897. Cited by: §4.6, §6.4. [8] A. Banerjee, L. Pasea, S. Harris, A. Gonzalez-Izquierdo, A. Torralbo, L. Shallcross, M. Noursadeghi, D. Pillay, N. Sebire, C. Holmes, et al. (2020) Estimating excess 1-year mortality associated with the COVID-19 pandemic according to underlying conditions and age: a population-based cohort study. The lancet 395 (10238), p. 1715â1725. Cited by: §4.6. [9] G. W. Brier et al. (1950) Verification of forecasts expressed in terms of probability. Monthly weather review 78 (1), p. 1â3. Cited by: §4.6.1. [10] British Heart Foundation (2025) CVD-COVID-UK / COVID-IMPACT. Note: https://bhfdatasciencecentre.org/areas/cvd-covid-uk-covid-impact/Accessed: 2025-01-10 Cited by: §4.8, §4.8. [11] British Heart Foundation (2025) Home- British Heart Foundation. Note: https://bhfdatasciencecentre.org/Accessed: 2025-07-20 Cited by: §4.8. [12] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. (2020) Language models are few-shot learners. Advances in neural information processing systems 33, p. 1877â1901. Cited by: §A.1, §3, §4.4. [13] Cabinet Office (2022) COVID-19 response: living with COVID-19. Technical report UK Government, London. Note: Accessed: 2026-05-26 External Links: Link Cited by: §4.6. [14] N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al. (2021) Extracting training data from large language models. In 30th USENIX security symposium (USENIX Security 21), p. 2633â2650. Cited by: §6.5. [15] J. Carrasco-Zanini, M. Pietzner, J. Davitte, P. Surendran, D. C. Croteau-Chonka, C. Robins, A. Torralbo, C. Tomlinson, F. GrĂźnschläger, N. Fitzpatrick, et al. (2024) Proteomic signatures improve risk prediction for common and rare diseases. Nature medicine 30 (9), p. 2489â2498. Cited by: §4.6.1, §6.4. [16] G. S. Collins, K. G. Moons, P. Dhiman, R. D. Riley, A. L. Beam, B. Van Calster, M. Ghassemi, X. Liu, J. B. Reitsma, M. Van Smeden, et al. (2024) TRIPOD+ ai statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 385. Cited by: Appendix C, §3. [17] T. Dao (2023) Flashattention-2: faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691. Cited by: §A.2. [18] Databricks (2024) Databricks Runtime 14.3 LTS. Note: https://docs.databricks.com/aws/en/release-notes/runtime/14.3ltsAccessed: 2025-04-29 Cited by: §4.8.2. [19] DMLC XGBoost (2022) XGBoost Documentation. Note: https://xgboost.readthedocs.io/en/stable/index.htmlAccessed: 2025-07-14 Cited by: §4.7. [20] A. Handy, A. Wood, C. Sudlow, C. Tomlinson, F. Kee, J. H. Thygesen, M. Mamouei, R. Sofat, R. Dobson, S. Ip, et al. (2021) A nationwide deep learning pipeline to predict stroke and COVID-19 death in atrial fibrillation. Medrxiv. Cited by: §3. [21] Health Data Research Gateway (2024) Civil Registration - Deaths. Note: https://healthdatagateway.org/en/dataset/877Accessed: 2025-08-13 Cited by: §4.1. [22] S. Hegselmann, G. von Arnim, T. Rheude, N. Kronenberg, D. Sontag, G. Hindricks, R. Eils, and B. Wild (2025) Large language models are powerful electronic health record encoders. arXiv preprint arXiv:2502.17403. Cited by: §6.2. [23] J. Hippisley-Cox, C. Coupland, and P. Brindle (2017) Development and validation of QRISK3 risk prediction algorithms to estimate future risk of cardiovascular disease: prospective cohort study. BMJ 357. Cited by: §3, §6.4. [24] Hugging Face (2023) Llama 2. Note: https://huggingface.co/docs/transformers/en/model_doc/llama2Accessed: 2025-04-28 Cited by: §A.2, §4.4. [25] A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira (2021) Perceiver: general perception with iterative attention. In International conference on machine learning, p. 4651â4664. Cited by: §6.1. [26] E. Jefferson, J. Liley, M. Malone, S. Reel, A. Crespi-Boixader, X. Kerasidou, F. Tava, A. McCarthy, R. Preen, A. Blanco-Justicia, et al. (2022) GRAIMATTER green paper: recommendations for disclosure control of trained machine learning (ml) models from trusted research environments (tres). arXiv preprint arXiv:2211.01656. Cited by: §6.5. [27] A. E. Johnson, L. Bulgarelli, L. Shen, A. Gayles, A. Shammout, S. Horng, T. J. Pollard, S. Hao, B. Moody, B. Gow, et al. (2023) MIMIC-iv, a freely accessible electronic health record dataset. Scientific data 10 (1), p. 1. Cited by: §6.4. [28] M. R. Junge and J. R. Dettori (2018) ROC solid: receiver operator characteristic (ROC) curves as a foundation for better diagnostic tests. Global spine journal 8 (4), p. 424â429. Cited by: §4.6.1. [29] R. Knight, V. Walker, S. Ip, J. A. Cooper, T. Bolton, S. Keene, R. Denholm, A. Akbari, H. Abbasizanjani, F. Torabi, et al. (2022) Association of COVID-19 with major arterial and venous thrombotic diseases: a population-wide cohort study of 48 million adults in england and wales. Circulation 146 (12), p. 892â906. Cited by: §3. [30] S. R. Knight, A. Ho, R. Pius, I. Buchan, G. Carson, T. M. Drake, J. Dunning, C. J. Fairfield, C. Gamble, C. A. Green, et al. (2020) Risk stratification of patients admitted to hospital with covid-19 using the isaric who clinical characterisation protocol: development and validation of the 4c mortality score. BMJ 370. Cited by: §6.4. [31] O. Kostopoulou, C. Tracey, and B. C. Delaney (2021) Can decision support combat incompleteness and bias in routine primary care data?. Journal of the American Medical Informatics Association 28 (7), p. 1461â1467. Cited by: §6.1. [32] Z. Kraljevic, D. Bean, A. Shek, R. Bendayan, H. Hemingway, J. A. Yeung, A. Deng, A. Baston, J. Ross, E. Idowu, et al. (2024) Foresightâa generative pretrained transformer for modelling of patient timelines using electronic health records: a retrospective modelling study. The Lancet Digital Health 6 (4), p. e281âe290. Cited by: §A.1, §A.3, §3, §4.4, §6.2. [33] Z. Kraljevic, J. A. Yeung, D. Bean, J. Teo, and R. J. Dobson (2024) Large language models for medical forecastingâforesight 2. arXiv preprint arXiv:2412.10848. Cited by: §6.2. [34] Y. Li, S. Rao, J. R. A. Solares, A. Hassaine, R. Ramakrishnan, D. Canoy, Y. Zhu, K. Rahimi, and G. Salimi-Khorshidi (2020) BEHRT: transformer for electronic health records. Scientific reports 10 (1), p. 7155. Cited by: §3, §3, §6.2. [35] P. J. Lillie, A. Samson, A. Li, K. Adams, R. Capstick, G. D. Barlow, N. Easom, E. Hamilton, P. J. Moss, A. Gow, et al. (2020) Novel coronavirus disease (Covid-19): the first two patients in the UK with person to person transmission. Journal of Infection 80 (5), p. 578â579. External Links: Document Cited by: §4.6. [36] G. Y. Lip, R. Nieuwlaat, R. Pisters, D. A. Lane, and H. J. Crijns (2010) Refining clinical risk stratification for predicting stroke and thromboembolism in atrial fibrillation using a novel risk factor-based approach: the euro heart survey on atrial fibrillation. Chest 137 (2), p. 263â272. Cited by: §3. [37] H. Logan Ellis, E. Palmer, J. T. Teo, M. Whyte, K. Rockwood, and Z. Ibrahim (2025) The early warning paradox. npj Digital Medicine 8 (1), p. 81. Cited by: §6.4. [38] Meta (2025) The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation. Note: https://ai.meta.com/blog/llama-4-multimodal-intelligence/Accessed: 2025-07-14 Cited by: §6.2. [39] K. G. Moons, J. A. Damen, T. Kaul, L. Hooft, C. A. Navarro, P. Dhiman, A. L. Beam, B. Van Calster, L. A. Celi, S. Denaxas, et al. (2025) PROBAST+ ai: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 388. Cited by: Appendix D, §3. [40] National Audit Office (2022) Managing NHS backlogs and waiting times in England. Note: https://w.nao.org.uk/reports/managing-nhs-backlogs-and-waiting-times-in-england/Accessed: 2026-05-25 Cited by: §4.6. [41] NHS Digital (2021) Data Provision Notice General Practice Data for Planning and Research. Note: https://w.spinneybrookmedcentre.co.uk/mf.ashx?ID=36ab53fe-7e31-4383-b4a-fd525cf8fa50Accessed: 2026-03-06 Cited by: §4.2. [42] NHS Digital (2022) Control of patient information (COPI) notice. Note: https://digital.nhs.uk/coronavirus/coronavirus-covid-19-response-information-governance-hub/control-of-patient-information-copi-noticeAccessed: 2025-07-14 Cited by: §3, §6.5. [43] NHS Digital (2023) COVID-19 Second Generation Surveillance System. Note: https://digital.nhs.uk/services/data-services-for-commissioners/datasets/covid-19-second-generation-surveillance-systemAccessed: 2025-01-10 Cited by: §4.1. [44] NHS Digital (2023) COVID-19 Vaccination Status. Note: https://digital.nhs.uk/services/data-services-for-commissioners/datasets/covid-19-vaccination-statusAccessed: 2025-01-10 Cited by: §4.1. [45] NHS Digital (2024) Data Access Request Service (DARS) products and services. Note: https://digital.nhs.uk/services/data-access-request-service-dars/dars-products-and-servicesAccessed: 2025-07-20 Cited by: §4.8. [46] NHS Digital (2025) Advisory Group for Data (AGD). Note: https://digital.nhs.uk/about-nhs-digital/corporate-information-and-documents/advisory-group-for-dataAccessed: 2025-07-20 Cited by: §4.8. [47] NHS Digital (2025) Clinical Classifications. Note: https://digital.nhs.uk/services/terminology-and-classifications/clinical-classificationsAccessed: 2025-04-14 Cited by: §4.2. [48] NHS Digital (2025) COVID-19 general practice extraction service (gpes) data for pandemic planning and research (GDPPR). Note: https://digital.nhs.uk/services/data-access-request-service-dars/dars-products-and-services/data-set-catalogue/gpes-data-for-pandemic-planning-and-research-gdpprAccessed: 2025-08-13 Cited by: §4.1. [49] NHS Digital (2025) Hospital Episode Statistics (HES). Note: https://digital.nhs.uk/data-and-information/data-tools-and-services/data-services/hospital-episode-statisticsAccessed: 2025-08-13 Cited by: §4.1. [50] NHS Digital (2025) SNOMED CT. Note: https://digital.nhs.uk/services/terminology-and-classifications/snomed-ctAccessed: 2025-04-14 Cited by: §4.2. [51] NHS England (2025) Secure Data Environment. Note: https://digital.nhs.uk/services/secure-data-environment-serviceAccessed: 2025-08-13 Cited by: §3, §4.8.1, §4.8.2, §4. [52] NVIDIA (2025) NVIDIA A10 Tensor Core GPU. Note: https://w.nvidia.com/en-gb/data-center/products/a10-gpu/Accessed: 2025-04-14 Cited by: §4.5, §4.8.2. [53] Office for National Statistics (2022) Coronavirus (COVID-19) Infection Survey, antibody data, UK: 18 May 2022. Note: https://w.ons.gov.uk/peoplepopulationandcommunity/healthandsocialcare/conditionsanddiseases/bulletins/coronaviruscovid19infectionsurveyantibodyandvaccinationdatafortheuk/18may2022Accessed: 2026-05-25 Cited by: §4.6. [54] B. Ozenne, F. Subtil, and D. Maucort-Boulch (2015) The precisionârecall curve overcame the optimism of the receiver operating characteristic curve in rare diseases. Journal of clinical epidemiology 68 (8), p. 855â859. Cited by: §4.6.1. [55] M. Pineda-MoncusĂ, F. Allery, H. Abbasizanjani, D. Powell, A. Prats-Uribe, J. H. Thygesen, A. Wood, C. Tomlinson, A. Banerjee, A. Akbari, et al. (2025) Ethnic disparities in COVID-19 mortality and cardiovascular disease in England and Wales between 2020â2022. Nature Communications 16 (1), p. 6059. Cited by: §6.1. [56] M. Pineda-MoncusĂ, F. Allery, A. Delmestri, T. Bolton, J. Nolan, J. H. Thygesen, A. Handy, A. Banerjee, S. Denaxas, C. Tomlinson, et al. (2024) Ethnicity data resource in population-wide health records: completeness, coverage and granularity of diversity. Scientific Data 11 (1), p. 221. Cited by: §3, §3, Figure 1, Figure 1, §4.1, Table 1, Table 1, §6.1. [57] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748â8763. Cited by: §6.1. [58] L. Rasmy, Y. Xiang, Z. Xie, C. Tao, and D. Zhi (2021) Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. NPJ digital medicine 4 (1), p. 86. Cited by: §3. [59] P. Renc, Y. Jia, A. E. Samir, J. Was, Q. Li, D. W. Bates, and A. Sitek (2024) Zero shot health trajectory prediction using transformer. NPJ Digital Medicine 7 (1), p. 256. Cited by: §A.1, §3, §3, §4.4, §4.5, §6.2, §6.2, §6.2, §6.3, §6.4. [60] R. D. Riley, G. S. Collins, L. Kirton, K. I. Snell, J. Ensor, R. Whittle, P. Dhiman, M. van Smeden, X. Liu, J. Alderman, et al. (2025) Uncertainty of risk estimates from clinical prediction models: rationale, challenges, and approaches. BMJ 388. Cited by: §4.5. [61] T. Saito and M. Rehmsmeier (2015) The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PloS one 10 (3), p. e0118432. Cited by: §4.6.1. [62] scikit-learn (2025) LogisticRegression. Note: https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegression.htmlAccessed: 2025-08-13 Cited by: §4.7. [63] A. Shmatko, A. W. Jung, K. Gaurav, S. Brunak, L. Mortensen, E. Birney, T. Fitzgerald, and M. Gerstung (2024) Learning the natural history of human disease with generative transformers. medRxiv. Cited by: §3, §6.2, §6.2, §6.2. [64] M. W. Sjoding, R. P. Dickson, T. J. Iwashyna, S. E. Gay, and T. S. Valley (2020) Racial bias in pulse oximetry measurement. New England Journal of Medicine 383 (25), p. 2477â2478. Cited by: §6.1. [65] E. Steinberg, J. Fries, Y. Xu, and N. Shah (2023) MOTOR: a time-to-event foundation model for structured medical records. arXiv preprint arXiv:2301.03150. Cited by: §6.2. [66] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu (2024) Roformer: enhanced transformer with rotary position embedding. Neurocomputing 568, p. 127063. Cited by: §A.2. [67] X. Su, S. Messica, Y. Huang, R. Johnson, L. Fesser, S. Gao, F. Sahneh, and M. Zitnik (2025) Multimodal medical code tokenizer. arXiv preprint arXiv:2502.04397. Cited by: §6.2. [68] The Lancet Rheumatology (2025) Translating ai innovation into clinical practice. The Lancet Rheumatology 7 (7), p. e451. Note: Published July 1, 2025 External Links: Document, Link, ISSN 2665-9913 Cited by: §6.1, §6.2. [69] J. H. Thygesen, C. Tomlinson, S. Hollings, M. A. Mizani, A. Handy, A. Akbari, A. Banerjee, J. Cooper, A. G. Lai, K. Li, et al. (2022) COVID-19 trajectories among 57 million adults in england: a cohort study using electronic health records. The Lancet Digital Health 4 (7), p. e542âe557. Cited by: §4.1, §4. [70] J. H. Thygesen, H. Zhang, H. Issa, J. Wu, T. Hama, A. Phiho-Gomes, T. Groza, S. Khalid, T. R. Lumbers, M. Hocaoglu, et al. (2025) Prevalence and demographics of 331 rare diseases and associated COVID-19-related mortality among 58 million individuals: a nationwide retrospective observational study. The Lancet Digital Health 7 (2), p. e145âe156. Cited by: §3, §6.1. [71] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. (2023) Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. Cited by: §A.2, §4.4. [72] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ĺ. Kaiser, and I. Polosukhin (2017) Attention is all you need. Advances in neural information processing systems 30. Cited by: §3. [73] S. Waxler, P. Blazek, D. White, D. Sneider, K. Chung, M. Nagarathnam, P. Williams, H. Voeller, K. Wong, M. Swanhorst, et al. (2025) Generative medical event models improve with scale. arXiv preprint arXiv:2508.12104. Cited by: §3, §3. [74] H. Wilde, C. Tomlinson, B. A. Mateen, D. Selby, H. K. Kanthimathinathan, P. Ramnarayan, P. Du Pre, M. Johnson, N. Pathan, A. Gonzalez-Izquierdo, et al. (2023) Hospital admissions linked to sars-cov-2 infection in children and adolescents: cohort study of 3.2 million first ascertained infections in england. BMJ 382. Cited by: Figure 6, Figure 6. [75] A. Wood, R. Denholm, S. Hollings, J. Cooper, S. Ip, V. Walker, S. Denaxas, A. Akbari, A. Banerjee, W. Whiteley, et al. (2021) Linked electronic health records for research on a nationwide cohort of more than 54 million people in england: data resource. BMJ 373. Cited by: §3, §4.1, §4. [76] World Health Organization (2025) International Statistical Classification of Diseases and Related Health Problems (ICD). Note: https://w.who.int/standards/classifications/classification-of-diseasesAccessed: 2025-04-14 Cited by: §4.2. [77] M. Wornow, S. Bedi, M. A. F. Hernandez, E. Steinberg, J. A. Fries, C. RĂŠ, S. Koyejo, and N. H. Shah (2024) Context clues: evaluating long context models for clinical prediction tasks on ehrs. arXiv preprint arXiv:2412.16178. Cited by: §6.1. [78] P. Wu, A. Gifford, X. Meng, X. Li, H. Campbell, T. Varley, J. Zhao, R. Carroll, L. Bastarache, J. C. Denny, et al. (2019) Mapping icd-10 and icd-10-cm codes to phecodes: workflow development and initial evaluation. JMIR medical informatics 7 (4), p. e14325. Cited by: §4.6, §5.2.2. [79] S. S. Zghebi, D. Reeves, C. Grigoroglou, B. McMillan, D. M. Ashcroft, R. Parisi, and E. Kontopantelis (2022) Clinical code usage in uk general practice: a cohort study exploring 18 conditions over 14 years. BMJ Open 12 (7), p. e051456. Cited by: §6.1. [80] S. Zhang, Q. Liu, N. Usuyama, C. Wong, T. Naumann, and H. Poon (2025) Exploring scaling laws for EHR foundation models. arXiv preprint arXiv:2505.22964. Cited by: §3, §6.2. [81] Y. Zhao, Y. Qu, K. Staniszewski, S. Tworkowski, W. Liu, P. MiĹoĹ, Y. Wu, and P. Minervini (2024) Analysing the impact of sequence composition on language model pre-training. arXiv preprint arXiv:2402.13991. Cited by: §6.2. [82] M. Zhu (2004) Recall, precision and average precision. Department of Statistics and Actuarial Science, University of Waterloo, Waterloo 2 (30), p. 6. Cited by: §4.6.1. Appendix A Training A.1 Objective Foresight-E was trained with an autoregressive next-token prediction objective, analogous to the paradigm used for LLMs [12, 32, 59]. Let N be the batch size, T the maximum sequence length, and V the vocabulary size. At each position t, given preceding tokens yn,<ty_n,<t, the model outputs a distribution over V. A binary mask mn,tm_n,t excludes padding and the initial demographic tokens from loss calculation. Excluding the initial tokens aims to mitigate the risk of biasing the model by learning to predict based solely on demographic features. The loss, âL, given the model weights, θ and the true token, yn,ty_n,t, is: â=â1ân,tmn,tân=1Nât=1Tmn,tlogP(yn,tâŁyn,<t,θ)L=- 1 _n,tm_n,t _n=1^N _t=1^Tm_n,t P(y_n,t y_n,<t,θ) (2) A.2 Architecture We adapted the Llama 2 transformer-decoder [71, 24] architecture with Rotary Positional Embeddings (ROPE) [66] and FlashAttention-2 [17]. Pretrained weights were not imported into the SDE due to governance restrictions and the custom vocabulary; therefore, training was conducted from scratch. We scaled the Llama decoder architecture down to 243 million parameters to facilitate efficient training on NVIDIA A10 GPUs. While we maintained the foundational hyperparameter ratios used across the Llama family [71], we adapted the final configuration to a 12-layer architecture with a hidden size of 1024, 8 attention heads, a feed-forward dimension of 4096, and a 1024-token context window. Additionally, input and output embeddings were tied to optimise parameter efficiency. A.3 Training Protocol We used bfloat16 mixed precision, attention dropout 0.1, gradient clipping (max norm 1.0), and weight decay 0.1. The Adam optimizer had β1=0.9 _1=0.9, β2=0.95 _2=0.95, with a linear warm-up over 3% of steps to a peak learning rate of 5Ă10â45Ă 10^-4, then cosine decay. The global batch size was 128, achieved via 8-way data parallelism and gradient accumulation (factor 2). Sequences were right-padded to 1024 tokens. The model was trained for one epoch, with validation loss evaluated every 1,000 steps on 32k sampled sequences. Training was completed in 4 days of wall-clock compute time distributed across 8 NVIDIA A10 GPUs. We did not conduct hyperparameter tuning due to resource constraints; instead, training parameters were selected based on standard defaults from prior literature [32, 3]. Appendix B Endpoint to Token Mapping Mapping of evaluation endpoints and associated figures to timeline tokens. An example phecode definition is shown. Endpoint Tokens Mortality DEATH Hospitalisation APC Emergency Hospitalisation APC_ADMIMETH_21, APC_ADMIMETH_22, APC_ADMIMETH_23, APC_ADMIMETH_24, APC_ADMIMETH_25, APC_ADMIMETH_28, APC_ADMIMETH_2A, APC_ADMIMETH_2B, APC_ADMIMETH_2C, APC_ADMIMETH_2D Phecode 8.0 - Intestinal Infection ICD10_A000, ICD10_A009, ICD10_A011, ICD10_A012, ICD10_A013, ICD10_A014, ICD10_A059, ICD10_A060, ICD10_A062, ICD10_A063, ICD10_A064, ICD10_A065, ICD10_A067, ICD10_A068, ICD10_A069, ICD10_A079 Appendix C TRIPOD+AI Checklist Assessment of Foresight-E against the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD)+AI Checklist [16]. The results section is excluded as not included in this paper. Item Checklist item Section Title Title Identify the study as developing or evaluating the performance of a multivariable prediction model, the target population, and the outcome to be predicted Title Abstract Title Identify the study as developing or evaluating the performance of a multivariable prediction model, the target population, and the outcome to be predicted Title Background Provide a brief explanation of the healthcare context and rationale for developing or evaluating the performance of all models Abstract Objectives Specify the study objectives, including whether the study describes model development, evaluation, or both Abstract Methods Describe the sources of data Abstract Describe the eligibility criteria and setting where the data were collected Abstract Specify the outcome to be predicted by the model, including time horizon of predictions in case of prognostic models Abstract Specify the type of model, a summary of the model-building steps, and the method for internal validation Abstract Specify the measures used to assess model performance (eg, discrimination, calibration, clinical utility) Abstract Results Report the number of participants and outcome events Pending Summarise the predictors in the final model Pending Report model performance estimates (with confidence intervals) Pending Discussion Give an overall interpretation of the main results Pending Registration Give the registration number and name of the registry or repository None Introduction Background Explain the healthcare context (including whether diagnostic or prognostic) and rationale for developing or evaluating the prediction model, including references to existing models 3 Describe the target population and the intended purpose of the prediction model in the context of the care pathway, including its intended users (eg, healthcare professionals, patients, public) 3 Describe any known health inequalities between sociodemographic groups 3 Objectives Specify the study objectives, including whether the study describes the development or validation of a prediction model (or both) 3 Methods Data Describe the sources of data separately for the development and evaluation datasets (eg, randomised trial, cohort, routine care or registry data), the rationale for using these data, and representativeness of the data 4.1 Specify the dates of the collected participant data, including start and end of participant accrual; and, if applicable, end of follow-up 4.1 Participants Specify key elements of the study setting (eg, primary care, secondary care, general population) including the number and location of centres 4.1 Describe the eligibility criteria for study participants 4.1 Give details of any treatments received, and how they were handled during model development or evaluation, if relevant 4.2 Data preparation Describe any data pre-processing and quality checking, including whether this was similar across relevant sociodemographic groups 4.1, 4.2 Outcome Clearly define the outcome that is being predicted and the time horizon, including how and when assessed, the rationale for choosing this outcome, and whether the method of outcome assessment is consistent across sociodemographic groups 4.6 If outcome assessment requires subjective interpretation, describe the qualifications and demographic characteristics of the outcome assessors N/A Report any actions to blind assessment of the outcome to be predicted None Predictors Describe the choice of initial predictors (eg, literature, previous models, all available predictors) and any pre-selection of predictors before model building 4.1, 4.2 Clearly define all predictors, including how and when they were measured (and any actions to blind assessment of predictors for the outcome and other predictors) 4.1, 4.3 If predictor measurement requires subjective interpretation, describe the qualifications and demographic characteristics of the predictor assessors N/A Sample size Explain how the study size was arrived at (separately for development and evaluation), and justify that the study size was sufficient to answer the research question. Include details of any sample size calculation 4.1 Missing data Describe how missing data were handled. Provide reasons for omitting any data 4.1, 4.3 Analytical methods Describe how the data were used (eg, for development and evaluation of model performance) in the analysis, including whether the data were partitioned, considering any sample size requirements 4.1 Depending on the type of model, describe how predictors were handled in the analyses (functional form, rescaling, transformation, or any standardisation) 4.1, 4.2, 4.3 Specify the type of model, rationaleâ , all model-building steps, including any hyperparameter tuning, and method for internal validation 4.4 Describe if and how any heterogeneity in estimates of model parameter values and model performance was handled and quantified across clusters (eg, hospitals, countries). See TRIPOD-Cluster for additional considerations None Specify all measures and plots used (and their rationale) to evaluate model performance (eg, discrimination, calibration, clinical utility) and, if relevant, to compare multiple models 4.6.1, 4.7 Specify all measures and plots used (and their rationale) to evaluate model performance (eg, discrimination, calibration, clinical utility) and, if relevant, to compare multiple models 4.6.1 Describe any model updating (eg, recalibration) arising from the model evaluation, either overall or for particular socio-demographic groups or settings None For model evaluation, describe how the model predictions were calculated (eg, formula, code, object, application programming interface) 4.5 Class imbalance If class imbalance methods were used, state why and how this was done, and any subsequent methods to recalibrate the model or the model predictions None Fairness Describe any approaches that were used to address model fairness and their rationale 4.1, A.1 Model output Specify the output of the prediction model (eg, probabilities, classification). Provide details and rationale for any classification and how the thresholds were identified 4.5, 4.6.1 Training versus evaluation Identify any differences between the development and evaluation data in healthcare setting, eligibility criteria, outcome, and predictors 4.1, 4.3,4.6.1 Ethical approval Name the institutional research board or ethics committee that approved the study and describe the participant informed consent or the ethics committee waiver of informed consent 4.8 Open Science Funding Give the source of funding and the role of the funders for the present study Blinded for review Conflicts of interest Declare any conflicts of interest and financial disclosures for all authors Blinded for review Protocol Indicate where the study protocol can be accessed or state that a protocol was not prepared Currently not publicly released Registration Provide registration information for the study, including register name and registration number, or state that the study was not registered Not registered Data sharing Provide details of the availability of the study data 4.8.1 Code sharing Provide details of the availability of the analytical code 4.8.2 Patient and public involvement Patient and public involvement Provide details of any patient and public involvement during the design, conduct, reporting, interpretation, or dissemination of the study or state no involvement 4.8 Discussion Interpretation Give an overall interpretation of the main results, including issues of fairness in the context of the objectives and previous studies Pending Limitations Discuss any limitations of the study (such as a non-representative sample, sample size, overfitting, missing data) and their effects on any biases, statistical uncertainty, and generalisability 6.1 Usability of the model in the context of current care Describe how poor quality or unavailable input data (eg, predictor values) should be assessed and handled when implementing the prediction model N/A Specify whether users will be required to interact in the handling of the input data or use of the model, and what level of expertise is required of users 6.5 Discuss any next steps for future research, with a specific view to applicability and generalisability of the model 6 Appendix D PROBAST+AI Assessment Assessment of Foresight-E against Risk Of Bias ASsessment Tool (PROBAST) + AI tool [39] Item Description Section/ Comment Step 1: PICOTS guidance Population Define the target population (e.g., patients) in whom the assessed prediction models are to be applied. The target population not only directs search strings and in/exclusion criteria of prediction models or prediction model studies in case of a systematic literature review, but also directs the applicability assessment. 4.1 Index Model Define the targeted prediction models to be assessed, which may be a single prediction model (the index model) of which the predictive accuracy is meta-analysed across multiple external evaluation studies of that index model but may also address multiple prediction models (developed or evaluated) for the targeted population, outcome or setting, depending on the assessorâs or prediction model review focus. 4.4 Comparator model(s) Define the other prediction models whose predictive ability is compared to that of the index model. 4.7 Outcome(s) Define the outcomes or endpoints that are predicted by the index (and possibly comparator) prediction models in the target population. 4.6 Timing Define the moment or time-point (e.g., in the patient work-up) at which the prediction with the prediction models is made (i.e., the start point or T0 of the use of the models). 4.6 Define the time or follow-up period in which the outcomes are being predicted by the prediction models in the targeted population (prediction horizon). 4.6 Setting and intended use of the prediction model Define the healthcare setting or context to which the index prediction models apply. The prediction ability of models may change across healthcare settings or contexts. 3 Step 2: Classify the type of prediction model assessment Development only Prediction model development only, i.e., without evaluation of its performance. â Evaluation only External validation of one or more existing models in new data Combination Prediction model development combined in the same study(publication) with the evaluation of its apparent performance, internal validation performance, or external validation performance. Step 3: Assess quality and applicability or risk of bias and applicability Participants and data sources Describe the sources of data and criteria for participant selection 4.1 Were appropriate data sources used? Yes Was an appropriate study design used? Yes Did the in- and exclusions of study participants result in a representative dataset? Yes Concern regarding quality of selection of participants and data sources Low Concern that the (data of the) included participants do not match the review question or the assessorâs intended use of the prediction model Low Predictors List and describe predictors included in the final prediction model, how they were defined and assessed, and their timing of assessment 4.1, 4.3 Were predictors defined and assessed in a similar way for all participants? Yes Was any pre-processing of predictors similar for all participants? Yes Were predictor assessments made without knowledge of outcome data? Yes Were the predictors included in the model available at the time the model was intended to be used? Yes Concern regarding the quality of the predictors or their assessment Low Concern that the definition, pre-processing, assessment, or timing of assessment of the predictors in the model do not match the review question or the assessorâs intended use Low Outcome Describe the outcome, how it was defined and determined, and the time interval between predictor assessment and outcome determination 4.6 At what time point was the outcome determined? If a composite outcome was used, describe the relative frequency/distribution of each contributing outcome? 4.6 Were outcomes defined and assessed appropriately? Yes Were outcomes defined and assessed in a similar way for all participants? Yes Were outcome assessments made without use or knowledge of predictor data? Yes Was the time interval between predictor assessment and outcome assessment appropriate? Yes Concern regarding quality of the outcome or its determination Low Concern that the outcome, its definition, assessment, or timing of assessment do not match the review question or the assessorâs intended use Low Analysis Describe the numbers of participants, number of candidate predictors, number of outcome events 4.1, 4.6 Describe how the prediction model was developed (e.g., with respect to modelling technique, predictor selection, and classification or risk group definition) 4.4 Describe the performance measures of the prediction model, e.g., (re)calibration, discrimination, (re)classification, net benefit, and whether they were adjusted for optimism 4.6.1 Describe missing data on predictors and outcomes as well as methods used for handling these missing data 4.3 Was there evidence that the sample size was reasonable? Yes Were continuous and categorical predictors handled appropriately? Yes Were participants with missing or censored data handled appropriately in the analysis? N/A If methods to address class imbalance were used, was the model or the model predictions recalibrated? N/A Were methods used to address potential model overfitting? Yes Concern regarding quality of the analysis Low Step 4: Assess the overall concerns regarding quality, risk of bias and applicability of the prediction model Overall concern regarding quality of the prediction model development Low concern regarding quality- If all four domains were rated low concern regarding quality. â High concern regarding quality- If at least one domain was rated high concern regarding quality. Unclear concern regarding quality- If at least one domain was rated unclear concern regarding quality and no domains were rated high concern. Overall concern regarding applicability of the prediction model development Low concern for applicability- If all three domains were rated low concern for applicability. â High concern for applicability- If at least one domain was rated high concern for applicability. Unclear concern for applicability- If at least one domain was rated unclear concern for applicability and no domains were rated high concern.