Paper deep dive
LLM-Augmented Computational Phenotyping of Long Covid
Jing Wang, Jie Shen, Amar Sra, Qiaomin Xie, Jeremy C Weiss
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 5:58:02 AM
Summary
The paper introduces 'Grace Cycle', an LLM-augmented computational phenotyping framework that iteratively refines hypotheses and extracts evidence from longitudinal patient data to identify clinical subphenotypes. Applied to 13,511 Long COVID participants from the NIH RECOVER initiative, the framework successfully categorized patients into 'Protected', 'Responder', and 'Refractory' phenotypes, demonstrating significant differences in symptom severity and vaccine response.
Entities (7)
Relation Signals (4)
Grace Cycle → analyzed → Long COVID
confidence 100% · In this work, we leverage large language models (LLMs) to analyze detailed clinical profiles of Long COVID participants
Grace Cycle → identified → Protected
confidence 95% · The framework identifies three distinct clinical phenotypes, Protected, Responder, and Refractory
Grace Cycle → identified → Responder
confidence 95% · The framework identifies three distinct clinical phenotypes, Protected, Responder, and Refractory
Grace Cycle → identified → Refractory
confidence 95% · The framework identifies three distinct clinical phenotypes, Protected, Responder, and Refractory
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Phenotypic characterization is essential for understanding heterogeneity in chronic diseases and for guiding personalized interventions. Long COVID, a complex and persistent condition, yet its clinical subphenotypes remain poorly understood. In this work, we propose an LLM-augmented computational phenotyping framework ``Grace Cycle'' that iteratively integrates hypothesis generation, evidence extraction, and feature refinement to discover clinically meaningful subgroups from longitudinal patient data. The framework identifies three distinct clinical phenotypes, Protected, Responder, and Refractory, based on 13,511 Long Covid participants. These phenotypes exhibit pronounced separation in peak symptom severity, baseline disease burden, and longitudinal dose-response patterns, with strong statistical support across multiple independent dimensions. This study illustrates how large language models can be integrated into a principled, statistically grounded pipeline for phenotypic screening from complex longitudinal data. Note that the proposed framework is disease-agnostic and offers a general approach for discovering clinically interpretable subphenotypes.
Tags
Links
- Source: https://arxiv.org/abs/2603.18115v1
- Canonical: https://arxiv.org/abs/2603.18115v1
Trouble viewing inline? Open PDF directly →
Full Text
44,694 characters extracted from source content.
Expand or collapse full text
Proceedings of Machine Learning Research LEAVE UNSET:1–12, 2026Submitted LEAVE UNSET; Published LEAVE UNSET LLM-Augmented Computational Phenotyping of Long Covid Jing Wangjing.wang20@nih.gov Jie Shenjie.shen@stevens.edu Amar Sraamarsra@email.gwu.edu Qiaomin Xieqiaomin.xie@wisc.edu Jeremy C Weissjeremy.weiss@nih.gov Abstract Phenotypic characterization is essential for un- derstanding heterogeneity in chronic diseases and for guiding personalized interventions. Long COVID, a complex and persistent con- dition, yet its clinical subphenotypes remain poorly understood. In this work, we propose an LLM-augmented computational phenotyp- ing framework “Grace Cycle” that iteratively integrates hypothesis generation, evidence ex- traction, and feature refinement to discover clinically meaningful subgroups from longitu- dinal patient data. The framework identifies three distinct clinical phenotypes, Protected, Responder, and Refractory, based on 13,511 Long Covid participants. These phenotypes ex- hibit pronounced separation in peak symptom severity, baseline disease burden, and longitu- dinal dose-response patterns, with strong sta- tistical support across multiple independent di- mensions. This study illustrates how large language models can be integrated into a principled, statistically grounded pipeline for phenotypic screening from complex longitudinal data. Note that the proposed framework is disease-agnostic and offers a general approach for discovering clinically interpretable subphenotypes. Data and Code Availability This data used for this study is part of the National Institutes of Health’s Researching COVID to Enhance Recovery (RECOVER) Initiative, which seeks to understand, treat, and prevent PASC https://recovercovid. org/, which is available by application. The under- lying code for this study is available as supplemental material. Hypothesis EvidenceLLM Figure 1: The pipeline of the LLM augmented phyenotyping for Long Covid, we called “Grace Circle”. We starts with a hypothe- sis, LLM looks for evidence from the data to support the hypothesis. With feedback about alignment with hypothesis and ev- idence by LLM, update the data and hy- pothesis. Repeat the process until the evi- dence are aligned with hypothesis. 1. Introduction Phenotying research is widely recognized as funda- mentally important in biology, medicine, and genet- ics. For example, diagnostic phenotyping is primarily used to support the diagnosis of disease Chen et al. (2022). By analyzing abnormal changes in behav- iors of physiological patterns, phenotyping can help identify early signs of a disease, such as sleep pat- tern changes leads to depression Kas et al. (2024). Predictive phenotyping aims to forecast future health events or risks of diseases, such as monitoring heart rate and physical activity for cardivascular diseases Pavarini et al. (2022). Monitoring phenotyping is used for ongoing monitoring of diagnosed health con- © 2026 J. Wang, J. Shen, A. Sra, Q. Xie & J.C. Weiss. arXiv:2603.18115v1 [cs.LG] 18 Mar 2026 Phenotyping based on LLM Table 1: Group comparisons across Protected, Responder, and Refractory cohorts. FeatureProtectedResponderRefractoryTest Statisticp-value Peak PASC Score4.98± 6.318.40± 7.122.80± 5.1 H = 4215.2 < 0.001 Initial Score4.988.5014.20F = 128.4 < 0.001 Avg. Vaccine Doses3.412.381.82F = 84.2 < 0.001 ditions or disease progression Langholm et al. (2023). This work studies belongs to monitoring phenotyp- ing. progression. We study the cohort with Long Covid from RECOVER program. Digital phenotyping is introduced in 2015 and of- fers new avenues for research and management of mental heath and physiological disease. It includes collecting health data digitally, supporting early dis- ease diagnosis and health management. The stream- ing data are from smart devices, sensors, and mo- bile apps about behavior, psychological, and phys- iological behavior Huckvale et al. (2019).Physi- ological phenotyping which involves monitoring an individual’s physiological paramters to assess their health. The physiological data includes breathing, sleep, heart rate, oxygen level, and activities. We find that breathing is most related to the PASC score. In this work, we use the wearable data of the par- ticipants in RECOVER program Zhang et al. (2025) which includes sleep, respiratory rate, heart rate, and physical activity. The global COVID-19 pandemic has last more than 4 years. More than 658 million people worldwide have been infected with SARS-CoV-2 (Long COVID) Or- ganization (2023). Long Covid can manifest in people across the life span, from children to older adults. It is a complex disease with sequelae across almost all or- gan systems. There are many subtypes that may have different risk factors (genetic, environmental, etc.) and distinct biologic mechanisms that may respond differently to treatments Al-Aly and Topol (2024). It is a major and ongoing public health challenge, with substantial impacts on quality of life and health-care utilization Danesh et al. (2023); Huang et al. (2021); Grady et al. (2025); Wang et al. (2025). The ob- servational analyses are very important. For exam- ple, it is suggested that use of the antiviral ritonavir- boosted nirmatrelvir within 5 days of symptom onset of SARS-CoV-2 infection may reduce the risk of Long Covid by 26% Xie et al. (2023). The diabetes drug metformin initiated within 7 days of SARSCoV-2 in- fection reduced the risk of Long Covid by 41% in a RCT Bramante et al. (2023). More evidence is needed to evaluate the effectiveness of reducing Long Covid risk and the safety of various antivirals. It is also important to understand different subtypes of Long Covid. Most recently, the deep learning tools has and AI tools have been used for healthcare related research, such as phenotyping. genetics, clinical trails Schmidt et al. (2026). In this work, we leverage large language models (LLMs) to analyze detailed clinical profiles of Long COVID participants enrolled in the NIH RE- COVER Program. Main result. We propose the pipeline to use LLM for auto clinical phenotyping as shown in Figure 3. Given the participants’ data with Long Covid, we first propose the weak assumption, then LLM reads the time series data of participants to seek evidence to support the hypothesis. If evidence does not sup- port hypothesis, we collect more features of the par- ticipants, revise the hypothesis, and let LLM find ev- idence again. The process converges when the evi- dences support the hypothesis. Then statistical anal- ysis is conducted to verify the significance of align- ment. It is related to the most recent works about LLM agent that can be updated automatically. How- ever, due to the complexity of the task in clinical domain, it requires human involvement for revising the hypothesis. It is a general pipeline. If we re- place the Long Covid data to another disease, then the final hypothesis would be new idea or discoveries that applies in that domain. For example, if the data becomes ICU patient discharge summaries, the hy- pothesis converges to the discoveries of ICU patient. Findings. In the analysis of data from 13,511 par- ticipants in the RECOVER adult cohort, a prospec- tive longitudinal cohort study. With PASC score as outcome, which is the evaluation of 44 symptoms (postexertional malaise, fatigue, brain fog, dizziness, gastrointestinal symptoms, palpitations, changes in sexual desire or capacity, loss of or change in smell or taste, thirst, chronic cough, chest pain, and abonor- mal movements), we set PASC less than 12 as State 0 and above 12 as State 1, as 12 is the suggested sever- ity threshold Thaweethai et al. (2023). We identify 2 Phenotyping based on LLM three clinical subpheynotyping based on the response of the vaccines. • Protected. 9,544 individuals (69% female; 8% Asia, 16% Hispanic/Latino; 17% non-Hispanic Black; median age, 44 years [IQR, 33-60]) main- taining Status 0 throughout the study. • Responder: 3,302 individuals (76% female; 6% Asia; 18% Hispanic/Latino; 14% non-Hispanic Black; median age, 49 years [IQR, 37-60]) who exhibited a state transition from Status 1 to Sta- tus 0. • Refractory. 665 individuals (82% female; 5% Asia; 16% Hispanic/Latino; 8% non-Hispanic Black; median age, 48 years [IQR, 39-57]) main- taining Status 1 regardless of intervention. Then we conduct the statistical analysis across multi- ple independent dimension, symptom severity, base- line status, treatment intensity to provide strong con- vergent validity for the discovered subphenotyping as shown in Table 1. • Peak PASC severity differs dramatically across subphenotypes.The Protected group shows a low mean peak PASC score, while the Re- sponder and especially the Refractory groups demonstrate substantially higher symptom bur- den. The extremely large Kruskal-Wallis statis- tic (H = 4215.2, p < 0.001) indicates that these differences not only statistically significant but also reflect a pronounced separation in symptom trajectories. • Baseline severity prior to vaccination increases monotonically from Protected to Responder to Refractory. the large F-statistic (F = 128.4, p < 0.001) supports the interpretation that the subphenotypes reflect differential vaccine- associated protection, rather than a uniform re- sponse across participants. The results indicate that the subphenotyping cap- tures clinically meaningful and biologically relevant structure in Long Covid, with direct implications for risk stratification, prognosis, and personalized inter- vention strategies. 2. Related works Phenotype research can be classified as the follow- ing categories based on applications scenarios, such as public health phenotyping and personalized health phenotyping Zhang et al. (2025). Our work belongs to public health phenotyping by aggregating and ana- lyzing the long term followup of a large number of in- dividuals, 13,511. Based on analysis objectives, phe- botyping research can be categorized as diagnostic phenotyping, predictive phenotyping, preventive phe- notyping and monitoring phenotyping. Our research belongs to the monitoring phenotyping. We con- tinuously monitor the participants with Long Covid their vaccine level, sleep, heart rate, lab test, aiding in Long Covid disease management. Based on the data source, phenotyping research can be mapped to behavioral phenotyping, physiological phenotyping, psychological phenotyping, environmental phenotyp- ing, social phenotyping, and medical phenotyping. This work mainly collects behavioral data, medical data, physiological data, and psychological data. The outcome in our work is PASC score which is based on 44 symptoms, such as brain fog, dizziness, chronic cough to name a few Thaweethai et al. (2023). Deep learning models, Large Language Models are powerful tools on phenotyping research and genomic measurements Avsec et al. (2026); Schmidt et al. (2024, 2026); Chen et al. (2025). For example, GPT- 4 has been used for sub-phenotyping of patients with Crohn’s disease, considering age at diagnosis and dis- ease behavior Schmidt et al. (2026). LLMs may offer an alternative to traditional bioinformatics methods to prioritize disease-associated genes based on disease phenotypes. Therefore, LLM based methods poten- tially enhance diagnostic accuracy and simplify the process for rare genetic diseases. GPT 3.5 has been used Peng et al. (2024). To protect our data are human related data, we use a local Large Language Model for our research. To solve complex tasks, prompt engineering has been demonstrated as an effective strategy , such as chain-of-thought prompting Wei et al. (2022); Brown et al. (2020). However, chain-of-thought prompt- ing requires a series of intermediate reasoning steps. In our case to explore the subphenotyping of Long Covid, there is not predefined steps. There are many works about reinforcement learning with human feed- back (RLHF) Ouyang et al. (2022); Bai et al. (2022); Zheng et al. (2023). In this work, we propose an auto feedback by comparing the evidence from LLM and hypothesis. Then we require update the hypothesis and data processing to feed LLM with more related data. Different from existing works that provide all information at once, we provide multiple iterations 3 Phenotyping based on LLM of data processing, hypothesis update and feedback collection. Pairwise comparison. Applications such as the dis- ease risk, treatment recommendation, it is impor- tant to consider the relationship between patients, like two patients with similar symptoms tend to re- cover with similar treatment Zhu et al.. In machine learning, pairwise comparisons have been used in preference learning in recommender systems, ranking and crowdsourced learning Xu and Davenport (2020); Gong et al. (2022); Zeng and Shen (2022a). This work focuses on learning threshold functions with pariwise comparisons Hopkins et al. (2020); Zeng and Shen (2022b). By adding weak distributional assumptions and allowing comparison queries, it makes the learn- ing algorithm requires exponentially fewer samples. In this work, we follow the line propose weak assump- tion first, then feed the similar trajectories of Long Covid to LLM to verify the hypothesis. PASC ScorePASC ScorePASC Score EnrollmentFollowup Va c c i n e Va c c i n e Va c c i n e Va c c i n e Followup Figure 2: The vaccine and wellness (PASC score) tra- jectory for Long Covid participants. 3. Method Formally, our data consists of several groups of fea- tures F i D i=1 , where F i = f 1 ,· ,f D i is the ith group of features with D i entities. For example, Fig- ure 2 shows the groups of vaccine record features and the PASC score outcome features. For vaccine record feature, each feature f i consists of two items e i ,t i where e i is the index of the vaccine, and e i is the recorded time of the event. For example, e i = 1st dose, t i =2022-02-20.For the target feature PASC score, each feature f i consists of two items e i ,t i where e i is the PASC score, and e i is the recorded datetime of the event. We also have wear- able features, For example, e i = weekly breathing rate, t i =2022-05-20. There are other features, such as weekly Breathing Rate and related summary date. Besides time series features, we also have statistic de- mographics features, such as sex and race. The goal of our project is to phenotyping of Long Covid to monitor the progression of Long Covid. The recovery of Long Covid is evaluated by the PASC score. To this end, we first need to find the features that are most related to the PASC score. Then based on that, we find the subphenotyping based on most related features. The end-to-end-pipeline is shown in Figure 3. We will introduce each component in the pipeline in the following section. Hypothesis Then we initialize a list assumptions H. For example, h 0 = does the Long Covid recover over time?, h 1 = does the Long Covid participants re- cover after Covid booster?, h 2 =Is Long Covid symp- toms related to breathing? The hypothesis is updated based on the LLM analysis of the features. Pairwise comparison Given the selected feature subset S = F 1 ,· ,F d . We compute the simi- larity between the participants in two ways. First, we compute the similarity based on the statistical analysis, such as the number of recorded vaccines / PASC Sore taken so far. Then we compute the se- mantic similarity by LLM based the embedding, such as MedGemma, Qwen3. Given the similarity between the participants, we feed the samples with certain fea- ture subset S individually, and with the top k similar pairs and in batches. Fairness Machine learning models have shown bi- ased predictions again disadvantaged groups on sev- eral real-world tasks Dressel and Farid (2018); Chai et al. (2025); Shen et al. (2022). Similar to accuracy, fairness can also be targeted by malicious adversari- als, leading to biased outcomes against certain demo- graphics. There are some techniques, including pre- processing that is to adjust training distribution to reduce discrimination Jiang and Nachum (2020); in- processing to impose fairness constraint during train- ing by reweighing or adding relaxed fairness regu- larization Jung et al. (2025); and post-processing to adjust the decision threshold in each sensitive group to achieve expected fairness parity Cruz and Hardt (2023). In this work, we use proprocessing to guaran- tee the fairness, first we filter demongraphic features in fairness comparison. The similarity threshold be- tween participants are computed with protected fea- tures. We also sample the data randomly multiple 4 Phenotyping based on LLM Hypothesis APASC PASC PASC B C Evidence LLM Pairwise comparison Fairness Ta b u l a r data Figure 3: The steps in the pipeline. LLM performs pairwise comparison, fairness criteria in the tabular data reading. The hypothesis updates by choosing different hypothesis from the pool. Evidence is updated by collecting related features, replacing or removing unrelated features from the knowledge base. The process continues until hypothesis and evidence are aligned. times to confirm the alignment between evidence and hypothesis. Group Feature Selection If LLM finds evidence in the given feature subset to support the hypothesis, then the process terminates, we conduct statistical analysis to verify the results. If the provided feature subset does not support the hypothesis, we have two actions. First is to revise the hypothesis based on the insights from LLM. The second is to update the fea- ture subset by selecting hypothesis related features and removing unrelated features. For example, the LLM discovers that weekly breathing rate is more re- lated to the PASC score, then in the next round, we update the hypothesis to be “how weekly breathing rate is related to PASC score over time?” The feature subset is updated by removing unrelated features, such as heart rate, activity, and include breathing related features, such as REM sleep breathing rate, all time/weekly/monthly breathing rate. LLM judgment Given a hypothesis h, we request certain a subset from F i D i=1 and find evidence to support the hypothesis. We use the thinking mode of the Large Language Models. If the features are aligned with the hypothesis, then LLM outputs the evidence from the features. If the features are not aligned with the hypothesis, LLM provides feature candidates that may be related to hypothesis. Figure 4 provides an example of the hypothesis and feature update process. The first hypothesis is assuming re- cover is related over time. But based on the LLM analysis result, there is no direct correlation. LLM suggests to include treatment, then we choose vac- cine as treatment. The updated hypothesis becomes the relationship between number of vaccines so far and recovery. Then the evidence supports the hy- pothesis. The iteration terminates. In practice, we could also include more features at first, and let LLM remove unrelated features. The example in Figure 4 includes low dimensional features first, which allows LLM to read as many participants’ records as possi- ble. Given the limit of the maximal lengths of input of LLM, there is a trade off between the number of records and the dimension of features. 4. Experiments 5. Dataset Our dataset is a subgroup of the RECOVER adult co- hort with adult participants enrolled before April 10, 2023. The analysis cohort included participants with 5 Phenotyping based on LLM Ta s k : Re a d t h e h i s t o r y o f t h e p a r t i c i p a n t S , PA S C score is the evaluation of the symptom, the lower the better. Find the evidence that support the hypothesis H , if the data does not support hypothesis, find most related features to PASC score. S_0: subject_id, 2024-01-20, PASC 10 subject_id, 2022-05-20, PASC 20 H_0: is the recover of Long Covid related to time? LLM_0: There is no relationship between data and the hypothesis, treatment maybe related with Long Covid recovery. S_1: subject_id, 2024-01-20, PASC 10, 1 dose subject_id, 2022-05-20, PASC 20, 2 doses H_1: is the recover of Long Covid related to the number of vaccines? LLM_1: The number of vaccines is highly related with Long Covid recovery. Figure 4: The update of hypothesis and feature space given the Long Covid cohort. The target variable is the PASC score which reflects the overall wellness of the participants of Long Covid. a study visit completed 6 months or more after the index date Horwitz et al. (2023). The participants are from 86 sites in 33 U.S. states, Washington, DC and puerto Rico, via facility and community-based out- reach. The participants complete quarterly question- naires about symptoms, social determinants, vaccina- tion stauts, and interim SARS-CoV-2 infections. In addition, participants contributes biospecimens and undergo physical and laboratory examinations at ap- proximately 0, 90 and 180 days from infection or neg- ative test date, and yearly thereafter. The primary outcome is onset of PASC, measured by signs and symptoms. The experiments are conducted on an AWS p3 in- stance with 8 GPUs. 5.1. LLM Judgment Given our longitudinal data, we allow the LLM to read and summarize the evidence directly. We use pairwise comparisons to jointly process multiple par- ticipants at a time (e.g., 4, 10, or 20 participants per batch). Through this process, the LLM identifies dis- tinct patterns: some participants consistently exhibit low PASC scores, some consistently high scores, and others show substantial fluctuations over time. Based on these observations, we examine the cor- relation between vaccination and recovery. As shown in Figure 5, individual vaccine events are not strongly associated with PASC scores. We then refine the analysis by introducing an additional feature, the cu- mulative number of vaccine doses. The LLM identi- fies dose 5 as a potential threshold beyond which a significant difference in PASC scores emerges. Finally, using the pipeline outputs on the selected samples, we construct a dataset consisting of time- series PASC scores and vaccination records, and con- duct downstream statistical analyses. 6. Findings with Statistical Analysis 6.1. Baselines we compared our approach against two standard methods for longitudinal data analysis:Linear Mixed-Effects Models (LMM) and Latent Class Tra- jectory Modeling (LCTM). We initially fit an LMM to evaluate general tem- poral trends and quantify population-level variance. The model yielded a fixed time effect of β = 0.010 and p-value=0.017 and a random group variance of σ 2 = 30.01. The magnitude of the group variance quantitatively confirms substantial between-subject heterogeneity within the dataset. This high degree of variance indicates that modeling the cohort via a single, population-averaged longitudinal trajectory is insufficient, necessitating the extraction of distinct patient subgroups. To derive data-driven patient subgroups, we ap- plied LCTM using the StepMix algorithm Morin et al. 6 Phenotyping based on LLM LLM: The ten participants cohort exhibits variable longitudinal wellness scores (PASC), with most individuals maintaining scores below the Long Covid threshold (>=12). Notably, patients who received multiple vaccine doses (Moderna/Pfizer) tended to show lower PASC scores over time compared to those with fewer vaccines or no clear vaccination history. For example, ** ** maintained consistently low PASC scores alongside multiple Moderna doses, while ** ** experienced higher scores before stabilizing post-vaccination. The strongest association on observed is the inverse relationship between cumulative vaccine doses and sustained low PASC scores. LLM: The four participants exhibit discount longitudinal wellness score trajectories. **** maintained consistently low PASC scores despite multiple Moderna vaccines and follow-ups with no intervening vaccines. **** showed progressive increases in PASC scores from 0 to 23, crossing the Long Covid threshold >12 by 2023-2024. ****demonstrated extreme variability, starting at 0-1, spiking to 16, and returning to 0, with Pfizer vaccines (doses 1-2) but no clear association between vaccination and score changes. **** had persistently high PASC despite Moderna vaccine (doses 1-3). The strongest observation was between prolonged elevated PASC and long Covid status, with no consistent evidence that vaccines improved outcomes across patient. LLM: **** PASC scores fluctuated between Long Covid and recovery, peaking at 12.0 in 2022-2023 but dropping to 3.0 in 2023 after Moderna vaccination between follow-ups. **** experience a significant decline from 12.0 at enrollment to 1.0 in 2023. **** maintained consistently low PASC score despite multiple Pfizer vaccines, suggesting robust protection against Long Covid. The strongest observed association link vaccine administration between follow-up visits to transient reductions in PASC scores. LLM: The findings align with hypothesis that vaccination is not a primary driver of Long Covid. **** and **** both vaccinated pre-enrollment, maintained consistently low PASC scores across follow-ups, indicating pre-infection vaccination may mitigate Long Covid severity. ****'s PASC scores showed peaks followed by declines, unrelated to post-enrollment vaccines. This implies PASC severity can fluctuate independently of vaccination after enrollment. **** received three doses pre-enrollment but still developed high PASC scores. Figure 5: Example of insights of LLM after reading trajectories of multiple participants. (2025), specifying a 3-class model. The resulting mix- ture model partitioned the patient trajectories as fol- lows: • Class 1 (29.7%): Characterized by a low initial baseline followed by a slightly decreasing trajec- tory over time. • Class 2 (37.9%): Exhibited a persistently high and stable trajectory throughout the observation window. • Class 3 (32.4%): Demonstrated a low initial baseline with a distinctly sharper subsequent de- cline compared to Class 1. While the morphological shapes of these latent classes broadly align with the phenotypes identified by our approach, the traditional LCTM baseline exhibited significant optimization limitations. Most notably, the StepMix model failed to achieve full conver- gence, and the covariance matrix for one of the ex- tracted classes collapsed, exhibiting near-zero vari- ance. This parametric instability suggests that tradi- tional trajectory clustering is ill-suited for the struc- tural complexities of this clinical setting. In contrast, our framework circumvents these convergence failures and class-collapse issues, yielding phenotypes that are both mathematically stable and highly interpretable. 6.2. Subphenotyping We discover 3 distinginshed Subphenotyping. The protected cohort with consistent low PASC score. The Responder with at least one status transfer from 1 to 0. The Refractory with consistent high PASC score. The demographic characteristics of three are shown in Table 2. Figure 6 (A) and (B) plots il- lustrates the distribution of peak and initial PASC scores across the three identified clinical cohorts: Pro- tected, Responder, and Refractory. 7 Phenotyping based on LLM Table 2: Demographic characteristics. CharacteristicProtectedResponder Refractory # Participants9,5443302665 Age, median (IQR)44.0 (33.0-60)49.0 (37-60)48.0 (39-57) Age category at enrollment 18-4532%41%48% 46-6452%44%43% >6516%15%9% Sex assigned at birth Female70%76%82% Male30%24%18% Race Asian8%6%5% Hispanic and Latino16%18%16% non-Hispani Black17%14%8% The Kruskal–Wallis test is performed at the par- ticipant level using peak PASC severity. For each par- ticipant, we extract the peak PASC score observed across all recorded time points. Each participant has a longitudinal record consisting of the observation date, the corresponding PASC score, and the cumu- lative number of vaccine doses received by that date. Participants are then assigned to one of three groups. Let n i denote the number of participants in group i, for i∈1, 2, 3, and let R i denote the sum of ranks of peak PASC scores in group i after pooling all partic- ipants. The Kruskal–Wallis H statistic is computed as H = 12 N (N + 1) 3 X i=1 R 2 i n i − 3(N + 1), where N = P 3 i=1 n i is the total number of partici- pants. Under the null hypothesis that the distribu- tions of peak PASC scores are identical across the three groups, H asymptotically follows a χ 2 distri- bution with 2 degrees of freedom. The effect size is 0.63. In our analysis, the test yields H = 4215.2 with a corresponding p-value < 0.001. It confirms that these cohorts represent distinct clinical trajectories of the disease. Stability of subphenotypes Using 100 bootstrap samples and Jaccard similarity, we get • Subphenotype Protected: 0.970 • Subphenotype Responder: 0.994 • Subphenotype Refractory: 0.972 All exceed the 0.85 threshold for strong stability Hen- nig (2007). 6.3. Dose Response Our analysis reveals that recovery is not a sponta- neous event but is closely tied to cumulative vaccine exposure, particularly within the “Responder” phe- notype. Following an initial post-immunization peak at Dose 1 (mean PASC = 10.38), Responders followed a consistent recovery trajectory with increasing cu- mulative doses (Figure 7). From Dose 2 through Dose 5, each additional dose corresponded to an av- erage reduction of ∼ 0.54 PASC points, reaching a 20.71% improvement by Dose 5 relative to the peak (P < 0.001; Dose 2: ≈ +7.31%; Dose 5: ≈ +20.71%). The 95% confidence band indicates stable estimation at common dose strata and widening uncertainty at sparsely populated extremes. In contrast, the protected cohort exhibited low baseline symptom burden with a shallow dose- response. PASC severity peaked at Dose 2 (mean = 1.79) and declined through Dose 5 (mean = 1.37), corresponding to a 23.17% improvement relative to the post-immunization peak (average reduction ≈ 0.14 points per additional dose from 2 to 5; P < 0.001). At higher cumulative doses, mean PASC con- tinued to trend downward (Dose 9: mean = 1.20, ≈ +32.67%; Dose 10: mean = 1.08, ≈ +39.70%), although estimates beyond Dose 8 were based on smaller sample sizes. The refractory cohort showed minimal early im- provement and a subsequent worsening pattern. Af- 8 Phenotyping based on LLM ProtectedResponderRefractory 0 5 10 15 20 25 30 Peak PASC score *** A ProtectedResponderRefractory 0 5 10 15 20 25 30 Initial PASC Score *** B Figure 6: Distribution of PASC severity across patient cohorts. A, Distribution of peak PASC severity and individual score variance. B, Distribution of initial PASC severity and individual score variance. 02468 Cumulative vaccine doses −10 0 10 20 30 40 50 Improvement vs peak (%) Peak severity Dose 1 (Mean=10.38) Protected Refractory Responder Figure 7: Longitudinal dose-response of PASC sever- ity differs by clinical phenotype. ter a modest peak at Dose 1 (mean = 20.54), PASC scores decreased to a shallow nadir at Dose 3 (mean = 19.63; ≈ +4.43% improvement; P = 0.011) but remained near ∼20 through Dose 6 (≈ +2.46%). At higher cumulative doses, symptom severity increased markedly (Dose 9: mean = 24.20, ≈ −17.84%; Dose 10: mean = 24.43, ≈ −18.95% relative to the Dose 1 peak), consistent with persistent and worsen- ing clinical burden. The consistency across the full cohort and phenotype-defined subgroups supports an interpreta- tion in which elapsed time alone does not correspond to spontaneous symptom resolution, while cumula- tive immunological exposure is associated with re- duced symptom burden in dose-sensitive phenotypes. 6.4. Time vs dose To separate passive temporal trends from immuno- logical exposure, we modeled PASC severity against (i) elapsed time since vaccination and (i) cumula- tive vaccine dose count. We have the conclusion that across all observations, severity increased modestly with time (r = 0.0521, P = 1.26× 10 −59 ), whereas cumulative vaccination showed an inverse association with severity (r = −0.0434, P = 5.95× 10 −42 ) (Fig- ure 8). Pearson correlation coefficients (r) between PASC severity and (i) time (Time-Severity) and (i) prior vaccine dose count across the full cohort and phenotype-defined subgroups are summarized in Fig- ure 8. Cell color encodes the direction and magnitude of r (diverging scale centered at 0), and each cell is annotated with the corresponding r value and statis- tical significance. Time-Severity correlations are pos- itive in all groups (All: r=0.052; Protected: 0.015; Responder: 0.020; Refractory: 0.048), indicating slightly higher severity with time. Dose-Severity cor- relations are negative in All, Protected, and Respon- der (All: r=-0.043; Protected: -0.037; Responder: - 0.036), consistent with a modest reduction in severity with higher prior vaccine dose count, while the Re- 9 Phenotyping based on LLM Time–SeverityDose–Severity All Protected Responder Refractory 0.052 *** -0.043 *** 0.015 *** -0.037 *** 0.020 *** -0.036 *** 0.048 ** 0.039 * PASC correlations with time and prior vaccination −0.050 −0.025 0.000 0.025 0.050 Correlation (r) Figure 8: PASC severity shows weak but significant correlations with time and prior vaccina- tion across phenotypes. fractory cohort shows a small positive Dose-Severity association (r=0.039).Significance is denoted as p < 0.05(∗), p < 0.01(∗), and p < 0.001(∗). 7. Conclusion We present a general, iterative pipeline “Grace Cy- cle” that leverages large language models (LLMs) for automated clinical phenotyping from longitudi- nal patient data.By treating hypothesis genera- tion, evidence extraction, and feature refinement as an interactive loop, the proposed framework enables LLMs to surface latent structure in complex clini- cal trajectories while remaining grounded in statis- tical validation. Applied to the RECOVER adult cohort, this approach identifies three clinically inter- pretable PASC subphenotypes—Protected, Respon- der, and Refractory, characterized by distinct symp- tom trajectories and differential responses to vacci- nation. The strong separation across peak symptom severity, baseline status, and treatment intensity pro- vides convergent evidence that these subphenotypes capture meaningful heterogeneity in Long COVID progression. More broadly, the framework is disease- agnostic and can be transferred to other clinical do- mains, where it may facilitate hypothesis discovery and accelerate data-driven clinical insights. 8. Limitations. This study has several limitations. First, although the LLM-guided pipeline automates evidence discov- ery and feature exploration, human expertise remains essential for hypothesis revision and clinical interpre- tation, particularly in high-stakes medical settings. Second, the analysis is observational in nature; there- fore, the identified associations between vaccination and PASC trajectories should not be interpreted as causal. Residual confounding factors, such as health- care access, comorbidities, or unmeasured behavioral variables, may influence observed outcomes. Third, PASC severity is derived from self-reported symp- toms, which are subject to recall bias and reporting variability. Finally, while the proposed framework is general, its performance and interpretability may de- pend on the quality and granularity of longitudinal data available in other disease domains. Future work will focus on incorporating causal modeling, improv- ing robustness to noisy clinical records, and reducing the need for human intervention in hypothesis refine- ment. References Ziyad Al-Aly and Eric Topol. Solving the puzzle of long covid. Science, 383(6685):830–832, 2024. ˇ Ziga Avsec, Natasha Latysheva, Jun Cheng, Guido Novati, Kyle R Taylor, Tom Ward, Clare Bycroft, Lauren Nicolaisen, Eirini Arvaniti, Joshua Pan, et al. Advancing regulatory variant effect predic- tion with alphagenome. Nature, 649(8099):1206– 1218, 2026. Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with re- inforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022. Carolyn T Bramante, John B Buse, David M Liebovitz,Jacinda M Nicklas,Michael A Puskarich, Ken Cohen, Hrishikesh K Belani, Blake J Anderson, Jared D Huling, Christopher J Tignanelli, et al. Outpatient treatment of covid-19 and incidence of post-covid-19 condition over 10 months (covid-out): a multicentre, randomised, quadruple-blind, parallel-group, phase 3 trial. The Lancet Infectious Diseases, 23(10):1119–1129, 2023. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, 10 Phenotyping based on LLM Arvind Neelakantan, Pranav Shyam, Girish Sas- try, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. Junyi Chai, Taeuk Jang, Jing Gao, and Xiaoqian Wang. On the alignment between fairness and ac- curacy: from the perspective of adversarial robust- ness. In ICML, 2025. I-Ming Chen, Yi-Ying Chen, Shih-Cheng Liao, and Yu-Hsuan Lin. Development of digital biomarkers of mental illness via mobile apps for personalized treatment and diagnosis. Journal of Personalized Medicine, 12(6):936, 2022. Qingyu Chen, Yan Hu, Xueqing Peng, Qianqian Xie, Qiao Jin, Aidan Gilson, Maxwell B Singer, Xuguang Ai, Po-Ting Lai, Zhizheng Wang, et al. Benchmarking large language models for biomed- ical natural language processing applications and recommendations. Nature communications, 16(1): 3280, 2025. Andr ́e F Cruz and Moritz Hardt.Unprocessing seven years of algorithmic fairness. arXiv preprint arXiv:2306.07261, 2023. Valerie Danesh, Alejandro C Arroliga, James A Bour- geois, Leanne M Boehm, Michael J McNeal, An- drew J Widmer, Tresa M McNeal, and Shelli R Kesler. Symptom clusters seen in adult covid-19 recovery clinic care seekers. Journal of general in- ternal medicine, 38(2):442–449, 2023. Julia Dressel and Hany Farid. The accuracy, fair- ness, and limits of predicting recidivism. Science advances, 4(1):eaao5580, 2018. Yu Gong, Greg Mori, and Fred Tung. Ranksim: Ranking similarity regularization for deep imbal- anced regression.In ICML, pages 7634–7649. PMLR, 2022. Connor B Grady, Bornali Bhattacharjee, Julio Silva, Jillian Jaycox, Lik Wee Lee, Valter Silva Monteiro, Mitsuaki Sawano, Daisy Massey, C ́esar Caraballo, Jeff R Gehlhausen, et al. Impact of covid-19 vac- cination on symptoms and immune phenotypes in vaccine-na ̈ıve individuals with long covid. Commu- nications Medicine, 5(1):163, 2025. Christian Hennig. Cluster-wise assessment of cluster stability. Computational Statistics & Data Analy- sis, 52(1):258–271, 2007. Max Hopkins, Daniel Kane, and Shachar Lovett. The power of comparisons for actively learning linear classifiers. Advances in Neural Information Pro- cessing Systems, 33:6342–6353, 2020. Leora I Horwitz, Tanayott Thaweethai, Shari B Bros- nahan, Mine S Cicek, Megan L Fitzgerald, Jason D Goldman, Rachel Hess, SL Hodder, Vanessa L Ja- coby, Michael R Jordan, et al. Researching covid to enhance recovery (recover) adult study protocol: Rationale, objectives, and design. Plos one, 18(6): e0286297, 2023. Chaolin Huang, Lixue Huang, Yeming Wang, Xia Li, Lili Ren, Xiaoying Gu, Liang Kang, Li Guo, Min Liu, Xing Zhou, et al. Retracted: 6-month con- sequences of covid-19 in patients discharged from hospital: a cohort study. The lancet, 397(10270): 220–232, 2021. Kit Huckvale, Svetha Venkatesh, and Helen Chris- tensen. Toward clinical digital phenotyping: a timely opportunity to consider purpose, quality, and safety. NPJ digital medicine, 2(1):88, 2019. Heinrich Jiang and Ofir Nachum. Identifying and correcting label bias in machine learning. In In- ternational conference on artificial intelligence and statistics, pages 702–712. PMLR, 2020. Hoin Jung, Junyi Chai, and Xiaoqian Wang. Adver- sarial latent feature augmentation for fairness. In The Thirteenth International Conference on Learn- ing Representations, 2025. Martien JH Kas, Niels Jongs, Maarten Mennes, Brenda WJH Penninx, Celso Arango, Nic van der Wee, Inge Winter-van Rossum, Jose Luis Ayuso- Mateos, Amy C Bilderbeck, Philippe l’Hostis, et al. Digital behavioural signatures reveal trans- diagnostic clusters of schizophrenia and alzheimer’s disease patients. European Neuropsychopharmacol- ogy, 78:3–12, 2024. Carsten Langholm, Andrew Jin Soo Byun, Janet Mullington, and John Torous. Monitoring sleep using smartphone data in a population of college students. Npj mental health research, 2(1):3, 2023. Sacha Morin, Robin Legault, F ́elix Lalibert ́e, Zsuzsa Bakk, Charles- ́ Edouard Gigu`ere, Roxane de la Sablonni`ere, and ́ Eric Lacourse. Stepmix: a python package for pseudo-likelihood estimation of gen- eralized mixture models with external variables. Journal of Statistical Software, 113:1–39, 2025. 11 Phenotyping based on LLM World Health Organization. Who coronavirus (covid- 19) dashboard. 2023. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730– 27744, 2022. Gabriela Pavarini, Aleksandra Yosifova, Keying Wang, Benjamin Wilcox, Nastja Tomat, Jessica Lorimer, Lasara Kariyawasam, Leya George, So- nia Al ́ı, and Ilina Singh. Data sharing in the age of predictive psychiatry: an adolescent perspective. BMJ Ment Health, 25(2):69–76, 2022. Di Peng, Liubin Zheng, Dan Liu, Cheng Han, Xin Wang, Yan Yang, Li Song, Miaoying Zhao, Yanfeng Wei, Jiayi Li, et al. Large-language models facili- tate discovery of the molecular signatures regulat- ing sleep and activity. Nature Communications, 15 (1):3685, 2024. Axel Schmidt, Magdalena Danyel, Kathrin Grund- mann, Theresa Brunet, Hannah Klinkhammer, Tzung-Chien Hsieh, Hartmut Engels, Sophia Pe- ters, Alexej Knaus, Shahida Moosa, et al. Next- generation phenotyping integrated in a national framework for patients with ultrarare disorders im- proves genetic diagnostics and yields new molecular findings. Nature genetics, 56(8):1644–1653, 2024. Linea Schmidt, Susanne Ibing, Florian Borchert, Ju- lian Hugo, Allison A Marshall, Jellyana Peraza, Judy H Cho, Erwin P B ̈ottinger, Bernhard Y Re- nard, and Ryan C Ungaro.Automating clini- cal phenotyping using natural language processing. Communications Medicine, 2026. Jie Shen, Nan Cui, and Jing Wang. Metric-fair active learning. In International conference on machine learning, pages 19809–19826. PMLR, 2022. Tanayott Thaweethai, Sarah E Jolley, Elizabeth W Karlson, Emily B Levitan, Bruce Levy, Grace A McComsey, Lisa McCorkell, Girish N Nadkarni, Sairam Parthasarathy, Upinder Singh, et al. Devel- opment of a definition of postacute sequelae of sars- cov-2 infection. Jama, 329(22):1934–1946, 2023. Jing Wang, Amar Sra, and Jeremy C Weiss. Ac- tive learning for forecasting severity among pa- tients with post acute sequelae of sars-cov-2. arXiv preprint arXiv:2506.22444, 2025. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural infor- mation processing systems, 35:24824–24837, 2022. Yan Xie, Taeyoung Choi, and Ziyad Al-Aly. Associa- tion of treatment with nirmatrelvir and the risk of post–covid-19 condition. JAMA Internal Medicine, 183(6):554–564, 2023. Austin Xu and Mark Davenport. Simultaneous pref- erence and metric learning from paired compar- isons. Advances in Neural Information Processing Systems, 33:454–465, 2020. Shiwei Zeng and Jie Shen. Efficient pac learning from the crowd with pairwise comparisons. In Inter- national Conference on Machine Learning, pages 25973–25993. PMLR, 2022a. Shiwei Zeng and Jie Shen. List-decodable sparse mean estimation. Advances in Neural Information Processing Systems, 35:24031–24045, 2022b. Yingbo Zhang, Jiao Wang, Hui Zong, Rajeev K Singla, Amin Ullah, Xingyun Liu, Rongrong Wu, Shumin Ren, and Bairong Shen.The compre- hensive clinical benefits of digital phenotyping: from broad adoption to full impact. NPJ Digital Medicine, 8(1):1–9, 2025. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595–46623, 2023. Dixian Zhu, Tianbao Yang, and Livnat Jerby. Gradi- ent aligned regression via pairwise losses. In ICML. 12