Paper deep dive
FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation
Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/20/2026, 4:17:15 AM
Summary
The paper introduces FairGlucose, a demographically balanced continuous glucose monitoring (CGM) benchmark comprising 300 patients across 12 strata (age, gender, diabetes type) to evaluate the equity of AI-based glucose forecasting. Benchmarking 33 models reveals that population-level validation metrics mask substantial subgroup disparities, particularly showing higher prediction errors for Type 1 Diabetes (T1D) patients compared to Type 2 Diabetes (T2D) patients. The study finds that specialized neural time-series models outperform frontier Large Language Models (LLMs), and behavioral event annotations contribute negligibly to prediction accuracy. The authors argue for subgroup-disaggregated reporting as a standard for digital health AI equity assessment.
Entities (8)
Relation Signals (6)
FairGlucose → contains → 300 patients
confidence 100% · FairGlucose comprises 300 patients sampled from a US-based mobile health platform
NS-Transformer → achievesbestperformance → FairGlucose
confidence 95% · NS-Transformer achieves the best performance (rMSE=25.6 mg/dL, Clarke A=76.4%)
Type 1 Diabetes → hashigherpredictionerror → Type 2 Diabetes
confidence 95% · T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001)
Population-level Validation → masks → Subgroup Disparities
confidence 95% · population-level external validation can conceal substantial subgroup disparities
Behavioral Events → contributesnegligiblyto → Prediction Accuracy
confidence 90% · behavioral events contribute negligibly (approximately 0.1 mg/dL) even under oracle event access
Frontier LLMs → underperforms → Specialized Neural Models
confidence 90% · Frontier LLMs underperform specialized neural models by 1–6 mg/dL
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows that subgroup performance gaps align with the proportion of clinically hard cases, and that input-length sensitivity varies across demographics, motivating personalized configurations. Frontier LLMs underperform specialized neural models by 1-6 mg/dL; behavioral events contribute negligibly (approximately 0.1 mg/dL) even under oracle event access. These findings establish that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard.
Tags
Links
- Source: https://arxiv.org/abs/2608.18296v1
- Canonical: https://arxiv.org/abs/2608.18296v1
Trouble viewing inline? Open PDF directly →
Full Text
133,836 characters extracted from source content.
Expand or collapse full text
FAIRGLUCOSE: A CGM FAIRNESS BENCHMARK REVEALS SUBGROUP DISPARITIES HIDDEN IN POPULATION-LEVEL VALIDATION A PREPRINT Junjie Luo ∗ Carey Business School Johns Hopkins University Baltimore, MD, USA Xuzhe Zhi Carey Business School Johns Hopkins University Baltimore, MD, USA Rui Han Welldoc, Inc. Columbia, MD, USA Abhimanyu Kumbara Welldoc, Inc. Columbia, MD, USA Anand K. Iyer Welldoc, Inc. Columbia, MD, USA Mansur E. Shomali Welldoc, Inc. Columbia, MD, USA Ritu Agarwal Carey Business School Johns Hopkins University Baltimore, MD, USA Guodong Gordon Gao Carey Business School Johns Hopkins University Baltimore, MD, USA ABSTRACT As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age×gender×type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medi- cation) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (≈1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows that subgroup perfor- mance gaps align with the proportion of clinically hard cases, and that input-length sensitivity varies across demographics, motivating personalized configurations. Frontier LLMs underperform specialized neural models by 1–6 mg/dL; behavioral events contribute negligibly (∼0.1 mg/dL) even under oracle event access. These findings establish that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard. Keywords continuous glucose monitoring·fairness·benchmark·time-series forecasting·digital health· diabetes Introduction Continuous glucose monitoring (CGM) is now embedded in a growing class of digital diabetes interventions, from real-time alerting and automated insulin delivery to mobile-app coaching, each ∗ Corresponding author: jluo41@jhu.edu arXiv:2608.18296v1 [cs.CY] 18 Aug 2026 relying on predictive models of near-term glucose trajectories [1–3]. As these AI-driven tools scale from research prototypes to clinical deployment, an unresolved question is whether their predictive accuracy is equitable across the demographic subgroups they will serve, or whether population-level performance metrics mask systematic subgroup-level failures [4–6]. This concern is not hypothetical. In clinical AI more broadly, recent work has shown that models with high average performance can underperform significantly for specific patient subgroups [7], and a scoping review of 467 AI fairness studies identified substantial evidence gaps across medical specialties [8]. In medical imaging, systematic demographic disparities have been documented [9]. Yet in CGM-based forecasting, fairness evaluation remains virtually absent. A major obstacle is the data itself. Existing CGM benchmarks are demographically imbalanced: GlucoBench [10] consolidates five public datasets totaling∼461 subjects but does not enforce demographic balance or include fairness metrics. OhioT1DM [11] provides only 12 type 1 diabetes patients, precluding generalization. GluFormer [12], a recent foundation model trained on 10 million CGM measurements, acknowledges that its training data is geographically and demographically limited but does not perform subgroup-level evaluation. Additionally, behavioral event annotations (meals, exercise, medication) are typically missing from these resources [13,14], preventing analysis of how glucose-altering events interact with prediction accuracy across demographics. The open question is therefore not whether CGM forecasting models are accurate on average, but whether population-level validation metrics provide a reliable guarantee of equitable performance across the diverse patient populations these tools will serve. To answer this question, we constructed FairGlucose, a CGM cohort specifically designed for fairness-aware evaluation: 300 patients balanced across 12 demographic strata (age×gender× type 1/type 2 diabetes), with 132,480 standardized forecasting samples, 3,945 unique timestamped behavioral events (meals, exercise, medication) logged by 81 of the 300 patients, and uniform data quality across all subgroups. Each sample contains a 24-hour CGM input and an 8-hour output window; the main benchmark reports the first 2 hours of that output window, with additional horizons reported in the Supplementary Information. Using this balanced infrastructure, we benchmarked 33 models spanning statistical, machine learning, neural time-series, and frontier LLM families on 2-hour glucose forecasting, with evaluation protocols that disaggregate performance by demographic subgroup [15]. We formulate two clinically meaningful evaluation tasks for CGM modeling: (1) prediction accuracy measured by root mean square error (rMSE) for 2-hour ahead glucose forecasting supporting real- time glycemic control, and (2) clinical safety assessed via Clarke Error Grid Analysis, reporting the percentage of predictions in the most clinically accurate zone (Clarke A), representing predictions with no clinical errors (within 20% or both<70 mg/dL) [16]. To support rigorous and equitable evaluation, we introduce a fairness-aware benchmarking protocol that integrates overall prediction error with subgroup-disaggregated performance reporting and event-aware metrics. We benchmark 33 models across four families: statistical methods (3), machine learning (4), neural time-series architectures (20, including the zero-shot foundation model Chronos2), and frontier LLMs and foundation models (6). Among neural models, attention-based architectures dominate the top ranks: NS-Transformer achieves the best performance (rMSE=25.6 mg/dL, Clarke A=76.4%), followed by Transformer (25.7), TimeXer (25.7), and PatchTST [17] (25.8). LightGBM [18] remains competitive as the best ML model (rMSE=26.0 mg/dL, Clarke A=76.8%), demonstrating that tree-based methods rival neural architectures for glucose forecasting. Foundation and LLM models show mixed results: TimeGPT [19] achieves competitive prediction error (rMSE=26.6), while general-purpose LLMs (GPT-5.1 [20], GPT-5 Mini [21], Claude 4.5 [22], Gemini 3 [23]) lag by 1–6 mg/dL, suggesting time-series-specific models adapt better for physiological forecasting. Baseline evaluation reveals key patterns. First, substantial subgroup disparities exist even with balanced data: T1D patients show consistently higher prediction error than T2D patients, and perfor- mance varies significantly across age-gender intersections. Although population-level generalization appears strong (OD/ID ratio≈1.0), subgroup-level ratios range from 0.8 to 1.4, demonstrating that aggregate metrics can mask heterogeneity. Second, input-length sensitivity is subgroup-dependent: older T2D patients require longer CGM histories for stable prediction, while younger T1D patients tolerate short input windows with minimal accuracy loss, supporting personalized input-length strate- gies rather than one-size-fits-all configurations. Third, instance-level difficulty analysis reveals that subgroup performance disparities align with the proportion of clinically hard cases (rapid glucose 2 Fair Glucose 18-39,T1D,Male 18-39,T1D,Female 18-39,T2D,Male 18-39,T2D,Female 40-64,T1D,Male 40-64,T1D,Female 40-64,T2D,Male 40-64,T2D,Female 65+,T1D,Male 65+,T1D,Female 65+,T2D,Male 65+,T2D,Female 20Held-in Patients 5Held- out Patients TrainValidationTestID Day1-15Day16-17Day18-20 TestOD Day1-12 OneQualifiedPatient-Day inputoutput Figure 1: FairGlucose Benchmark Curation Process. 300 patients stratified across 12 balanced subgroups (age: 18–39, 40–64,≥65×gender: M/F×diabetes type: T1D/T2D), 25 per subgroup. Each subgroup has 20 held-in patients (train/valid/test-id: 15/2/3 days) and 5 held-out patients (test-od: 12 days). Each day yields 24 prediction pairs (24h input + 8h output). Includes 3,945 unique event annotations (meals, exercise, medication) logged by 81 patients across 132,480 prediction pairs. excursions, post-meal spikes); top models concentrate gains on easy cases yet offer minimal improve- ment on hard cases, precisely the scenarios where accurate prediction is most clinically valuable. Fourth, explicit event features provide small but consistent improvements: incorporating glucose- altering events (meals, exercise, medication) as model features yields minor performance gains (∼0.1 mg/dL) even under oracle event access (including post-observation events), with benefits consistent across demographic subgroups. These findings demonstrate that equitable glucose prediction requires careful attention to subgroup composition, input configuration, difficulty-aware evaluation, and event handling. Our work makes three contributions. (1) Clinical finding. Population-level external-validation metrics (≈1.0) systematically mask sub- stantial subgroup-level variation (0.8–1.4 range) in CGM forecasting accuracy. This masking persists across all 33 evaluated models spanning four paradigms, suggesting that the disparity is a task- level property rather than an architecture-level artifact. Subgroup performance gaps align with the proportion of clinically hard cases, and input-length sensitivity varies by demographic, motivating personalized configurations. Together these findings demonstrate that population-level validation alone is insufficient for equity assessment of digital health AI tools. (2) Resource. FairGlucose is a 300-patient demographically balanced CGM cohort with 132,480 standardized forecasting samples and 3,945 unique timestamped behavioral events logged by 81 patients, to be released with held-in/held-out splits, an evaluation server, and a 33-model reference leaderboard, under a data-use agreement. (3) Model benchmarking. Frontier LLMs (GPT-5.1, Claude 4.5, Gemini 3) underperform specialized neural models by 1–6 mg/dL, while behavioral events contribute negligibly (∼0.1 mg/dL) to top models even under oracle event access, suggesting modern forecasters implicitly capture event-driven dynamics from CGM signal alone. Results Study design and dataset overview. FairGlucose comprises 300 patients sampled from a US-based mobile health platform, evenly stratified across 12 demographic subgroups defined by age (18–39 / 40–64 /≥65), gender (male / female), and diabetes type (T1D / T2D), with 25 patients per subgroup (Figure 1; see Methods for full cohort and protocol details). Each patient contributes multiple 24-hour CGM traces sampled at 5-minute intervals; from these we extract 132,480 input–output prediction pairs comprising a 24-hour input window and an 8-hour output window, alongside 3,945 unique timestamped behavioral events (meals, exercise, medication) logged by 81 patients. The main benchmark evaluates the first 2 hours of the output window. We benchmark 33 baseline 3 Table 1: Glycemic characteristics of the FairGlucose cohort by demographic stratum (all 25 patients per stratum, held-in and held-out pooled; split-wise breakdowns in Supplementary Note A). Each stratum contains 25 patients (20 held-in for training/validation/test-ID; 5 held-out for test-OD). Glucose values in mg/dL; CV is the coefficient of variation; TIR is time-in-range (70–180 mg/dL); MAGE is the mean amplitude of glycemic excursions. StratumNMean glucoseCVTIR (%)MAGE Overall and marginal subgroups Overall300162.70.2667.14.06 Age 18–39100163.80.2665.94.40 Age 40–64100165.90.2664.94.23 Age≥65100158.50.2570.63.57 Female 150162.90.2567.94.11 Male150162.60.2666.34.02 Type 1 (T1D)150163.70.2965.34.20 Type 2 (T2D)150161.80.2368.93.92 12-way intersectional strata (age× gender× type) 18–39 F T1D25165.70.2964.24.67 18–39 F T2D25160.60.2269.04.02 18–39 M T1D25174.60.3059.04.67 18–39 M T2D25154.10.2371.34.24 40–64 F T1D25163.10.2765.04.12 40–64 F T2D 25169.30.2366.14.32 40–64 M T1D 25161.00.3065.54.32 40–64 M T2D 25170.40.2262.94.15 65+ F T1D25157.50.2772.33.89 65+ F T2D25161.00.2270.93.62 65+ M T1D25160.50.2965.83.55 65+ M T2D 25155.10.2473.23.20 Table 2: Dataset partition overview. Each cell shows the per-subgroup value (total across 12 subgroups in parentheses). SplitPatientsPatient-DaysI/O Pairs Held-in Train20 (240)300 (3,600)7,200 (86,400) Valid20 (240)40 (480)960 (11,520) Test-ID20 (240)60 (720)1,440 (17,280) Held-outTest-OD5 (60)60 (720)1,440 (17,280) Total25 (300)460 (5,520)11,040 (132,480) models spanning four families: statistical (Naïve, ARIMA, AutoETS), machine learning (LightGBM, XGBoost, CatBoost, Linear), neural time-series (20 architectures including PatchTST, iTransformer, TimeXer, Crossformer, TFT, N-HiTS, N-BEATS, DLinear, TimeMixer, TimesNet, Chronos2, and others; see Methods), and frontier LLMs and foundation models (TimeGPT, GPT-5.1, GPT-5 Mini, Claude 4.5 Sonnet/Haiku, Gemini 3 Flash), on 2-hour-ahead forecasting. Predictive accuracy is reported as root-mean-square error (rMSE) and Clarke Error Grid Zone A percentage (Clarke A); equity is quantified as cross-subgroup disparity in rMSE. Glycemic characteristics of the cohort. Glycemic profiles differ measurably across the 12 demo- graphic strata even after balancing for sample size (Table 1). T1D patients exhibit higher coefficient of variation (CV∼0.29) and lower time-in-range (TIR∼65%) than T2D patients (CV∼0.23, TIR ∼69%), consistent with the clinical literature on subtype-level glycemic variability. Within-subtype, younger T1D males show the highest variability (mean glucose 174.6 mg/dL, CV 0.30). These differences motivate subgroup-disaggregated evaluation and provide the substrate against which model performance is later compared. Behavioral event annotations. A defining feature of FairGlucose is comprehensive timestamped behavioral events that accompany the CGM signal. Of the 300 patients, 81 (27.0%) logged at least one event via the mobile health platform, producing 3,945 unique behavioral events: medication events (2,809; 71.2%), meal events with nutritional details (648; 16.4%), and exercise events with duration and intensity (488; 12.4%). Logging patients recorded a median of 38 events (mean 48.7, range 1–237) over their observation period. Because the hourly sliding-window design produces overlapping 48-hour observation windows, each unique event appears in approximately 30 prediction 4 Model rMSE (mg/dL)Clarke A (%) AllT1DT2DFemaleMale18–3940–6465+AllT1DT2DFemaleMale18–3940–6465+ ARIMA27.4430.7824.1027.8926.9927.7628.0526.5174.6770.2679.0774.1475.1974.1174.5575.38 Naive29.0332.4225.6329.5428.5129.4129.7827.8572.8668.6177.1172.4573.2771.9872.9073.76 AutoETS29.7233.3526.0930.4429.0030.2730.5328.3172.9668.6377.2972.4173.5171.5373.0374.46 LightGBM25.9629.0722.8426.3425.5826.5926.5824.6376.77*72.4581.10*76.30*77.24*75.25*77.16*78.08* Linear26.0428.9223.1526.5125.5626.8026.5024.7575.5771.6479.5074.9776.1773.8776.2276.74 XGBoost 27.8130.9824.6428.4127.2028.0728.7426.5675.1570.6279.6874.5775.7374.1874.9476.49 CatBoost28.4631.5625.3629.0627.8729.1029.0727.1973.8169.9477.6873.2374.3872.3774.1275.03 NS-Transformer25.62*28.46*22.7926.02*25.23*25.9926.39*24.4476.4572.70*80.2075.8677.0375.2276.6177.61 Transformer25.6528.5622.74*26.0225.2826.1026.4624.34*76.4372.4580.4175.8777.0075.1676.2478.02 TimeXer25.7528.5322.9626.1225.3825.97*26.4724.7676.0672.0780.0575.3276.8074.9376.3377.02 PatchTST25.7528.6822.8326.0825.4325.9826.5724.6675.9872.0279.9375.2076.7575.0775.9976.97 Crossformer 25.8528.7023.0026.2925.4126.0626.6024.8675.7971.8679.7375.0576.5374.6276.0276.82 Informer25.9428.7923.0926.2925.5926.1226.6525.0176.1372.1880.0975.5476.7375.0076.4577.05 TFT26.0228.7823.2526.4025.6426.2126.7525.0675.5271.6879.3774.8576.2074.7275.7176.20 MICN26.1128.9623.2626.4525.7726.3626.8825.0475.5771.6679.4875.0276.1274.4775.8176.50 TimesNet26.4029.4523.3526.7526.0526.6827.0125.4675.6171.2979.9275.1176.1074.3376.1576.45 DLinear26.4029.2923.5226.6526.1626.6327.0325.5175.1571.0479.2674.7475.5774.0775.6275.83 iTransformer26.4929.4923.5026.7426.2526.7827.1925.4475.0770.9979.1574.5075.6473.7275.3576.29 SCINet26.5029.3223.6826.8326.1726.8027.0925.5675.1271.2079.0374.7075.5373.9275.5975.91 N-HiTS26.7830.0723.4926.9626.6027.0927.3025.8974.7070.2479.1574.2875.1273.5575.0875.55 TiDE26.8029.7123.8827.0526.5426.9327.4226.0174.5070.4378.5874.0774.9473.6574.8075.13 TSMixer26.8029.8523.7527.0426.5526.9427.5325.8774.2669.9778.5573.8074.7273.3374.6474.87 TimeMixer26.9329.9423.9227.2126.6527.2127.6425.8974.3470.0978.5973.9874.7073.0874.7275.31 N-BEATS27.4230.5724.2827.6527.2027.6528.0026.6173.3368.8577.8172.9873.6772.3173.7873.93 FEDformer31.1734.6727.6731.5830.7631.0231.6730.8266.9861.7772.1966.3867.5866.8567.3466.78 Autoformer35.6339.7131.5636.3434.9334.8536.0336.1260.6854.5066.8559.7361.6261.3961.7858.72 Chronos239.7044.2135.1840.1439.2538.8639.7140.6354.2247.9260.5253.9254.5255.2555.8451.38 TimeGPT26.6329.7423.5327.2925.9826.9627.3925.5275.5571.4079.7074.7276.3874.7975.3276.62 GPT-5.127.5730.6224.5228.0327.1027.8928.8325.9074.9170.6079.2274.3975.4374.0774.2376.58 GPT-5 Mini27.6730.7824.5728.2627.0927.9128.8426.2374.8170.4879.1574.0875.5473.9174.3476.31 Claude 4.5 Sonnet28.2030.8725.5328.8727.5428.5029.7126.3073.1769.6476.7072.2274.1171.7972.6775.20 Claude 4.5 Haiku28.7731.7825.7629.3128.2329.1529.8427.2673.3369.1977.4772.5174.1671.9573.3574.87 Gemini 3 Flash31.2934.1528.4231.8830.6931.2732.6929.8569.9265.6374.2169.1370.7269.4369.4071.08 Table 3: Model performance on rMSE (mg/dL) and Clarke Error Grid Zone A (%) at the 2-hour prediction horizon, averaged across test-ID and test-OD splits, across all subgroups defined by diabetes type (T1D, T2D), gender (Female, Male), and age (18–39, 40–64, 65+). Clarke A represents clinically accurate predictions. Higher Clarke A percentage indicates better clinical accuracy.Underlined values indicate the best performance within each model family. Bolded values highlight the top-2 overall performers per subgroup: * denotes the best performance and bold without asterisk denotes the second best. pairs; consequently 22,239 of 132,480 samples (16.8%) contain at least one event within their window. Event density varies systematically across demographic subgroups: patients aged 65+ log 22% fewer events than younger groups, while T2D patients log 12% more events than T1D, reflecting differences in app engagement and disease-management style. Event timing is highly asymmetric: 33.8% of event occurrences fall in the 24-hour input window, 3.0% in the evaluated 2-hour forecast window, and 63.2% in the 2–24 hours after the prediction time. Event-aware results should therefore be interpreted as retrospective event-aware analyses or upper-bound estimates unless restricted to events available at prediction time. Overall Results Model Family Performance. Table 3 shows neural time-series models achieve the best per- formance. Attention-based architectures dominate the top ranks: NS-Transformer ranks first (rMSE=25.6 mg/dL, Clarke A=76.4%), followed closely by Transformer (25.7), TimeXer (25.7), and PatchTST (25.8). LightGBM achieves competitive performance as the best ML model (rMSE=26.0 mg/dL, Clarke A=76.8%), demonstrating that tree-based methods remain viable al- ternatives to neural architectures for glucose forecasting. ML models achieve competitive performance, with LightGBM within 0.2 mg/dL of PatchTST on prediction error (25.96 vs 25.75) and in fact ahead of it on clinical accuracy (Clarke A 76.8% vs 76.0%). ARIMA remains competitive as a statistical baseline, demonstrating that classical time series methods provide strong performance for well-structured physiological data. Foundation and LLM models show mixed performance. TimeGPT achieves competitive prediction error (rMSE=26.6 mg/dL, within 1.0 mg/dL of top neural performers), while Chronos2 as a zero- shot foundation model lags substantially (rMSE=39.7). Frontier LLMs (GPT-5.1, GPT-5 Mini, Claude 4.5, Gemini 3) show larger gaps (rMSE: 1–6 mg/dL behind top neural models; Clarke A: 70–75%), suggesting that general-purpose LLMs are not yet competitive with specialized time-series architectures for physiological forecasting [5]. Clinical Accuracy Profile. Clarke Error Grid analysis shows that top-performing models achieve strong clinical accuracy, with LightGBM reaching 76.8% of predictions in Zone A (clinically accurate), followed by NS-Transformer and Transformer (both 76.4%), and Informer and TimeXer 5 24 26 28 30 32 34 36 38 40 2h_rMSE (Mean +/- 95% CI) 2h_rMSE by Model (test-id & test-od) StatsMLNeuralAPI test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Autoformer DLinear N-HiTS TFT PatchTST NS-Trans. Gemini 3 Claude 4.5 GPT-5.1 TimeGPT 0.975 1.000 1.025 1.050 OD / ID Ratio Figure 2: Generalization performance across models. Top: 2-hour rMSE on test-ID and test-OD with 95% CIs. Bottom: OD/ID ratio. Figure created by the authors using Matplotlib. T1DT2D 0 5 10 15 20 25 30 35 40 TestID 2h_rMSE Age: 18-39 T1DT2D 0 5 10 15 20 25 30 35 40 TestID 2h_rMSE Age: 40-64 T1DT2D 0 5 10 15 20 25 30 35 40 TestID 2h_rMSE Age: 65+ Mean 2h_rMSE Male Female 0.6 0.8 1.0 1.2 1.4 OD / ID Ratio MFMF T1DT2D 0.6 0.8 1.0 1.2 1.4 OD / ID Ratio MFMF T1DT2D 0.6 0.8 1.0 1.2 1.4 OD / ID Ratio MFMF T1DT2D Figure 3: Intersectional subgroup accuracy and generalization across 12 demographic strata. Figure created by the authors using Matplotlib. (76.1%) (Table 3). The remaining 23–24% of predictions fall into zones B through E, with the majority in Zone B (benign errors requiring no treatment adjustment). Across the 20 neural architectures, Clarke A rates range from 54.2% (Chronos2, zero-shot) to 76.4% (NS-Transformer), demonstrating that architecture selection substantially impacts clinical accuracy. Performance Across Subgroups. Subgroup patterns are consistent across models. T1D patients show consistently higher errors than T2D patients (rMSE difference 6.3 mg/dL; 95% bootstrap CI [4.3, 8.2]; permutationp < 0.001), reflecting the greater glucose variability and insulin dynamics characteristic of type 1 diabetes. This gap persists on held-out patients (test-od:∆ = 5.6mg/dL, 95% CI [1.0, 9.6],p = 0.011). T2D patients also achieve notably higher clinical accuracy (Clarke A: 80.1% vs 72.2% for T1D among top-5 models), further demonstrating the prediction challenges posed by type 1 diabetes. At the marginal level, gender differences are small and not significant (∆ M−F =−0.19mg/dL, 95% CI[−2.12, 1.76],p = 0.86), and younger adults (18–39) show higher 6 error than the oldest group without reaching significance (∆ 18–39 vs 65+ = 1.85mg/dL, 95% CI [−0.63, 4.66],p = 0.16). This patient-level marginal test should not be read against the per-model gender columns of Table 3, which are sample-weighted and show a consistent female-minus-male gap near+0.8mg/dL; the marginal difference is small precisely because the gender effect changes sign across diabetes types. Our balanced sampling protocol confirms these disparities reflect true physiological heterogeneity rather than sample size artifacts. Generalization Across Distributions. Figure 2 shows overlapping confidence intervals between test-id and test-od across all models, with OD/ID ratios near 1.0, indicating robust generalization to out-of-distribution patients. While population-level generalization appears strong, subgroup-level analysis reveals hidden heterogeneity in both accuracy and generalization behavior, motivating deeper fairness evaluation. Fairness Analysis Intersectional Subgroup Disparities. Figure 3 reveals intersectional patterns across age, gender, and diabetes type. The gender pattern is conditional on diabetes type rather than uniform: among T1D patients males show lower error than females at ages 18–39 (∆ F−M = +3.3mg/dL) and 40–64 (+2.1), while among T2D patients females show lower error at 40–64 (−3.4) and 65+ (−2.4), and the T1D male advantage has disappeared by age 65+ (−0.5). Critically, OD/ID ratios range from 0.8 to 1.4 across subgroups despite the population-level average near 1.0, demonstrating that aggregate metrics mask subgroup vulnerabilities. For example, T2D males aged 65+ generalize robustly (OD/ID =0.90), while T2D females aged 40–64 degrade substantially (OD/ID=1.42), the widest gap in the cohort. These findings underscore the necessity of intersectional fairness evaluation rather than single-axis demographic analysis. Accuracy-Fairness Tradeoffs. To assess both accuracy and fairness jointly, Figure 4 plots each model’s mean 2-hour rMSE (x-axis) against subgroup performance disparity (y-axis), measured as rMSE standard deviation across the 12 demographic strata. Models in the bottom-left quadrant (low error, low disparity) are preferred for equitable deployment. Neural models PatchTST and TFT achieve this ideal combination, demonstrating that state-of-the-art accuracy does not require sacrificing fairness. Gradient-boosted models sit mid-range on disparity (LightGBM 6.23, XGBoost 6.34) despite LightGBM’s competitive mean error, and they are not the least equitable family: the neural model N-HiTS (6.58) is worse than both. The statistical baselines are in fact the least equitable of the reasonable performers: ARIMA (6.68) and Naïve (6.79) rank 13th and 14th of 15 models on disparity, with Autoformer the outlier on both axes (error 35.6 mg/dL, disparity 8.14). The frontier LLMs separate on the two axes rather than clustering: they carry high mean error (Gemini 3 31.3 mg/dL, Claude 4.5 28.2 mg/dL) yet some of the lowest cross-subgroup disparity in the panel (Claude 4.5 at 5.34 is the lowest of all 15 models; Gemini 3 5.73 is third lowest). Low disparity here reflects uniformly mediocre zero-shot prediction rather than equitable competence, so it is not evidence of deployment readiness; it does show that error magnitude and error equity are separable properties that a single leaderboard column would hide. TimeGPT, a time-series-specific foundation model, sits closer to the neural models than to the general-purpose LLMs on mean error (26.6 mg/dL) with mid-range disparity (6.21). Input Length Sensitivity Subgroup-Level Patterns. To assess whether longer input histories benefit all patient subgroups equally, we computed the temporal ratio (TR), defined as the rMSE at a given input lengthLdivided by the rMSE at the longest available input (288 steps, 24 hours), for each subgroup (see Methods for definition). A TR near 1.0 indicates that shorter inputs preserve full performance; higher values indicate degradation. Figure 5 (left) reveals that the relative utility of input length varies substantially across demographic and clinical subgroups. Among younger patients (ages 18–39), particularly those with T1D, shorter input windows (7 or 37 steps) yield performance close to the 288-step baseline, suggesting that recent glucose history is sufficient for accurate forecasting in these groups. In contrast, older patients (ages 65+), especially those with T2D, experience larger performance degradation when input length is reduced, implying greater dependence on longer historical context. Gender-specific trends also emerge: male and female patients within the same age and disease category respond differently to 7 262830323436 Mean Error (lower is better) 5.5 6.0 6.5 7.0 7.5 8.0 Fairness Metric (lower is better) 2h_rMSE Mean vs Spread Naive ARIMA LightGBM CatBoost XGBoost PatchTST TFT N-HiTS DLinear Autoformer TimeGPT GPT-5.1 GPT-5 Mini Claude 4.5 Gemini 3 Model Family llm ml stats neural Figure 4: Model accuracy vs. subgroup disparity. Models near bottom-left (low error, low disparity) are preferred. PatchTST and TFT achieve strong performance on both axes. Figure created by the authors using Matplotlib. input window truncation, indicating complex interactions between physiology, disease progression, and temporal information. These findings demonstrate that the optimal input history length is not uniform across populations, highlighting the potential for subgroup-adaptive input-length strategies. Model-Level Patterns. Figure 5 (right) investigates how varying input lengths affect model per- formance across families. Neural models, particularly PatchTST, TFT, and N-HiTS, exhibit the largest gains as input length increases, especially from short (7 or 37 steps) to long (145 or 288 steps) windows, indicating their capacity to capture long-range temporal dependencies in glucose dynamics. In contrast, linear models show marginal improvement, and foundation models (TimeGPT, LLM-based) show modest gains, suggesting limitations in their zero-shot temporal modeling capabil- ities. These findings challenge prior claims that extended CGM input length provides limited benefit [24,25], showing instead that the benefit is model-dependent: architectures equipped to leverage longer histories gain substantially, with implications for both model selection and personalized input configuration. Instance-Level Difficulty Difficulty Distribution Across Subgroups. To investigate model robustness across prediction samples of varying difficulty, we define high-difficulty (hard) and low-difficulty (easy) instances using a majority-vote approach across models (see Methods for protocol). Instances where most models produce large errors are labeled hard; those where most produce small errors are labeled easy. 8 1.00 1.05 1.10 1.15 1.20 1.25 1.30 1.35 Performance Ratios T1DT2D Age: 18-39 7 37 73 145 7 37 73 145 7 37 73 145 7 37 73 145 MFMF T1DT2D Age: 40-64 7 37 73 145 7 37 73 145 7 37 73 145 7 37 73 145 MFMF T1DT2D Age: 65+ 7 37 73 145 7 37 73 145 7 37 73 145 7 37 73 145 MFMF (a) Subgroup-level sensitivity 7 3773 145289 7 3773 145289 7 3773 145289 7 3773 145289 7 3773 145289 7 3773 145289 7 3773 145289 25 30 35 40 2h_rMSE Arima PatchTST TFT N-HiTS DLinear Chronos2 N-BEATS (b) Model-level sensitivity Figure 5: Input-length sensitivity analysis. (a) Subgroup-level temporal ratio (TR) across 12 de- mographic strata: shorter input windows degrade performance more for older T2D patients. (b) Model-level 2-hour rMSE across input lengths: neural architectures benefit most from extended input windows. Figure created by the authors using Matplotlib. 0.0 0.2 0.4 0.6 0.8 1.0 Proportion MFMFMFMFMFMF T1DT2DT1DT2DT1DT2D Age: 18-39Age: 40-64Age: 65+ 12.4 23.9 52.9 12.5 25.6 55.1 11.1 23.1 49.4 11.6 23.1 51.1 12.8 24.5 52.5 11.9 24.2 56.4 12.1 23.8 52.9 11.4 23.6 51.5 12.3 24.4 56.6 13.3 24.3 53.9 11.3 22.8 51.4 11.6 23.1 50.9 Easy Neutral Hard (a) Difficulty distribution by subgroup 0.70.80.91.01.11.21.3 Hard / Naive Hard 0.7 0.8 0.9 1.0 1.1 1.2 1.3 Easy / Naive Easy Off-chart: Autoformer (1.61) Hard vs Easy Performance Stats ML Neural API Naive Naive ARIMA AutoETS LightGBM CatBoost XGBoost NS-Trans. PatchTST TFT TimeXer N-HiTS DLinear TimeGPT GPT-5.1 Claude 4.5 Best Hard Ratio: PatchTST Best Easy Ratio: TimeXer (b) Hard vs. easy performance by model Figure 6: Instance-level difficulty analysis. (a) Difficulty distribution across 12 demographic strata: hard cases concentrate in T1D subgroups. (b) Hard vs. easy case performance by model (normalized by Naive): top models concentrate gains on easy cases with limited improvement on hard cases. Figure created by the authors using Matplotlib. Figure 6 (left) shows that the distribution of hard and easy cases varies substantially across demo- graphic strata. Young male T1D and middle-aged female T1D patients have the highest proportion of hard cases, while older T2D males show the highest proportion of easy cases. These patterns indicate that prediction difficulty reflects underlying physiological and behavioral complexity rather than random variation. While the share of hard and easy cases differs across subgroups, the average rMSE within each difficulty level is relatively stable: hard cases cluster around 63 mg/dL and easy cases around 10 mg/dL. This consistency suggests that instance-level difficulty, not subgroup identity alone, primarily drives prediction error. Subgroup disparities in overall performance therefore largely reflect differences in the proportion of hard cases. Hard cases often correspond to clinically critical periods such as rapid glucose excursions, post-meal spikes, or exercise-induced drops, scenarios where accurate prediction is most valuable. Evaluation metrics focused on average performance may underweight these high-risk moments, masking potential real-world performance gaps. Model Robustness Across Difficulty. Figure 6 (right) compares model performance on easy and hard instances, normalized by the Naive baseline. Top neural models improve on both easy cases (easy ratios 0.77–0.89) and hard cases (hard ratios 0.89–0.92), but the gains are asymmetric: models achieve larger relative improvements on easy cases than on hard cases. PatchTST achieves the best 9 hard ratio (0.89), while TimeXer achieves the best easy ratio (0.77). This asymmetry reveals that top-performing models concentrate their aggregate accuracy gains on easy cases, with more limited improvement on hard cases, precisely the clinically critical scenarios where accurate prediction is most needed. Statistical baselines (ARIMA, AutoETS) maintain ratios near 1.0 on both easy and hard cases, while frontier LLMs show variable performance. Future work should prioritize difficulty-stratified evaluation and hard-case-targeted training to improve both clinical robustness and fairness. Beyond demographic fairness, input-length sensitivity, and instance difficulty, we examine how behavioral events (meals, exercise, medication) impact prediction difficulty and whether event handling varies equitably across patient subgroups, adding a temporal-behavioral dimension to fairness evaluation. Table 4: Comparison of event vs noevent model performance at 2-hour prediction horizon across subgroups. Event features include both historical (pre-observation) and future (post-observation) events; these results represent an upper-bound estimate of prospective event-aware forecasting.∆= Event - NoEvent: Negative values (dark green,∆ <−0.1) indicate event model is better, positive values (dark red,∆ ≥ 0.1) indicate noevent model is better. Differences with|∆| < 0.1mg/dL are shown in black (negligible). Note that the models in this table are the Nixtla/neuralforecast implementations used for the event ablation, whereas Table 3 reports the TSLib implementations; VanillaTrans. here is therefore not the same code path as Transformer there, and the two should not be read across. Only the Event vs. NoEvent contrast within a column is meaningful. Subgroup Event rMSE (mg/dL)NoEvent rMSE (mg/dL)∆ (Event - NoEvent) PatchTSTTFTAutoformerVanillaTrans.PatchTSTTFTAutoformerVanillaTrans.PatchTSTTFTAutoformerVanillaTrans. All25.3726.0037.9547.5225.4826.0838.7147.72-0.11-0.08-0.77-0.20 T1D28.2128.9442.3048.6628.3628.9743.0448.90-0.15-0.04-0.74-0.25 T2D22.5323.0633.5946.3722.6023.1834.3946.53-0.07-0.12-0.80-0.16 Female 25.7826.3938.5549.0725.9126.5139.2649.24-0.13-0.13-0.71-0.16 Male24.9625.6137.3445.9625.0525.6438.1746.20-0.09-0.03-0.83-0.24 18–3925.6926.2337.0847.3525.9226.3337.8347.36-0.23-0.10-0.74-0.00 40–6426.1026.7438.1446.9326.3326.7938.9247.28-0.23-0.05-0.77-0.35 65+24.2824.9938.6748.4024.1225.0739.4648.65+0.16-0.08-0.79-0.25 Event Impact on Performance Event Prevalence and Demographics. Event annotations enable stratified evaluation of prediction difficulty under glucose-altering conditions. Among the 16.8% of samples containing events (22,239 of 132,480 total), medication events are most prevalent (72.8% of window-level event occurrences), followed by diet (15.9%) and exercise (11.3%). Event density varies across demographics: patients aged 65+ log 22% fewer events than younger groups, while T2D patients log 12% more events than T1D patients, reflecting differences in disease management patterns and app engagement. Event-Stratified Prediction Difficulty. To quantify how event features affect prediction accuracy, we compare models with event features enabled (“Event”) versus disabled (“NoEvent”) (Table 4). Because event features include post-observation events not necessarily available at prediction time, these results represent an upper-bound estimate of prospective event-aware forecasting performance. Table 4 reveals a surprising finding: events barely affect accuracy for top-performing models. PatchTST and TFT show differences of only−0.11and−0.08mg/dL respectively (negative values indicating event model performs slightly better), far below clinical significance thresholds (∼5 mg/dL). This pattern holds across all subgroups: T1D patients show−0.15to−0.25mg/dL differ- ences, age groups span+0.16to−0.35mg/dL, and both genders exhibit differences near−0.10 mg/dL. In contrast, Autoformer shows larger event-related improvements (−0.77mg/dL), while Vanilla- Transformer shows minimal difference (−0.20mg/dL). Across all 32 comparisons (4 models×8 subgroups), models with event features consistently match or outperform those without; in fact, 31 of 32 show the event model performing better or equivalently. This suggests that explicit event features provide small but consistent gains, though modern neural forecasters may implicitly learn event-driven glucose dynamics from 24-hour temporal patterns even without explicit event annotations. Event-Related Fairness Analysis. To assess fairness, we examine whether event features benefit all demographic subgroups equally. Table 4 shows consistent patterns across 32 model-subgroup comparisons: 22 show event models performing better (∆ < −0.1mg/dL), 9 show negligible 10 differences (|∆| < 0.1), and only 1 shows the noevent model slightly better (PatchTST for 65+: ∆ = +0.16 mg/dL). The consistency across subgroups indicates event features do not create performance disparities. For top-performing models (PatchTST, TFT), event benefits typically range from−0.08to−0.23 mg/dL for most subgroups; one exception (PatchTST for 65+) shows+0.16mg/dL, though this difference remains clinically negligible. T1D and T2D patients show similar patterns, males and females benefit comparably, and age groups show only minor variation. This uniformity suggests that the evaluated neural forecasters handle available behavioral event features similarly across diverse patient profiles, supporting further evaluation of event-aware forecasting without immediate evidence of demographic-specific event-feature harm. Discussion Our principal finding is that population-level external-validation metrics can create false reassurance about CGM forecasting equity. Aggregate out-of-distribution performance appears stable (≈1.0), yet subgroup-level ratios range from 0.8 to 1.4 across 12 demographic strata. This gap is not limited to any single model: it persists across all 33 evaluated models spanning four paradigms (statistical, machine-learning, neural time-series, and frontier LLM/foundation), suggesting that the disparity reflects the prediction task itself. Because the FairGlucose cohort is demographically balanced by construction, these disparities cannot be attributed to unequal sample sizes. This has direct implications for how CGM-based AI tools should be qualified before clinical deploy- ment. If evaluation relies solely on population-level metrics, subgroups such as young T1D females, who carry the highest in-distribution error in this cohort (31.9 mg/dL against 19.9 mg/dL for the best-served subgroup), may be systematically underserved; notably the worst-generalizing subgroup is a different one again (T2D females aged 40–64, OD/ID=1.42), so accuracy and generalization single out different groups and both need reporting. Subbaswamy et al. [7] demonstrated this problem in general clinical AI; our results show that analogous concerns arise in CGM forecasting across 33 models. Recent reviews of AI fairness in healthcare [8,9] have identified evidence gaps across medical specialties; FairGlucose addresses this gap for CGM-based digital diabetes tools. We recommend that subgroup-disaggregated reporting become a default standard for digital health AI evaluation, analogous to how clinical trials report outcomes by demographic subgroup. Compared to existing CGM benchmarks, FairGlucose addresses several key limitations. GlucoBench [10] aggregates five public datasets (∼461 subjects) but does not enforce demographic balance, mak- ing subgroup-level fairness evaluation unreliable. OhioT1DM [11] provides only 12 type 1 diabetes patients. GluFormer [12], trained on 10 million CGM measurements, acknowledges demographic limitations but does not perform subgroup evaluation. FairGlucose provides 300 patients across 12 balanced strata with behavioral event annotations, enabling a systematic fairness evaluation of CGM forecasting models. Cohort metadata, splits, code, and the reference leaderboard will be made available athttps://github.com/JHU-CDHAI/FairGlucoseupon acceptance; raw CGM data can be requested under a data-use agreement with Welldoc, Inc. Two additional dimensions reinforce this conclusion. Input-length sensitivity varies across subgroups: older T2D patients require longer CGM histories for stable prediction, while younger T1D patients tolerate short input windows with minimal accuracy loss. This motivates personalized or subgroup- adaptive input-length strategies rather than fixed-length configurations. Instance-level difficulty analysis further shows that subgroup performance disparities align with the proportion of clinically hard cases (rapid glucose excursions, post-meal spikes, exercise-induced drops). Top-performing neural and ML models concentrate their aggregate accuracy gains on easy cases, yet offer minimal improvement on these clinically critical hard cases compared to simple baselines like ARIMA. This asymmetry suggests that average-case evaluation metrics can underweight high-risk scenarios, and that difficulty-stratified reporting should accompany subgroup-disaggregated evaluation for comprehensive equity assessment. Several important limitations warrant attention. First, event annotations are limited to app-logged data (meals, exercise, medication); passive sensing modalities (sleep, stress, physical activity patterns) are not captured. Second, data originates from a US-based mobile health platform; generalization to other populations, geographic regions, and CGM devices should be validated. Third, the 12 subgroup stratification (age, gender, diabetes type) does not capture all fairness-relevant factors such 11 as race, socioeconomic status, or comorbidities. Fourth, there is risk of overfitting to specific dataset properties; we encourage model development on multiple datasets alongside FairGlucose. Future directions for the community include: (1) expanding to multi-site cohorts with broader demo- graphic and device coverage; (2) developing event-aware and personalized forecasting approaches; (3) investigating fairness-aware model architectures using the subgroup performance metrics; and (4) incorporating hybrid or mechanistic models for interpretability. The standardized evaluation protocols can inform regulatory approval and clinical deployment decisions. FairGlucose provides the infrastructure needed for reproducible, fairness-focused CGM forecasting development. FairGlucose provides a demographically balanced CGM benchmark with behavioral event annotations and a 33-model reference leaderboard, enabling reproducible fairness research in glucose forecasting. We hope this infrastructure accelerates development of fairness-aware CGM forecasting and supports more equitable diabetes care. Methods This section details the construction of the FairGlucose benchmark and the protocol used for baseline evaluation. We describe (i) cohort construction and dataset partitioning, (i) input-output sample con- struction and event annotation, (i) the predictive-accuracy and fairness metrics used for evaluation, and (iv) the baseline models and training setup. Dataset Partitioning The dataset is constructed from de-identified CGM records collected via a mobile health platform. Data were collected from existing users of a US-based mobile health digital platform for diabetes man- agement, reflecting naturalistic self-management behaviors rather than controlled study conditions. Participants used either Dexcom G6 or Abbott FreeStyle Libre continuous glucose monitors, both factory-calibrated devices providing readings every 5 minutes with manufacturer-stated accuracy of approximately 9–10% mean absolute relative difference (MARD). The mixed-device cohort ensures model robustness across CGM platforms, reflecting real-world deployment scenarios where patients may use different devices. Incomplete CGM traces are excluded via the qualified patient-day criterion described below; no imputation is applied. As shown in Figure 1, to capture physiological and behavioral heterogeneity, we stratify patients across three axes: age (18–39, 40–64,≥65), gender (male, female), and diabetes type (Type 1, Type 2), yielding 12 disjoint subgroups. No additional personal information beyond the reported stratification variables is used for modeling or evaluation. We sample 25 patients from each subgroup, ensuring sufficient CGM records for robust evaluation. Each subgroup is partitioned into 20 held-in and 5 held-out patients. Hold-in patients contribute to training, validation, and in-distribution testing. Specifically, we extract 20 qualified days per held-in patient and assign 15 days to training set, the next 2 to validation set, and final 3 days to the in-distribution (test-id) set. Held-out patients are entirely excluded from model development and each provides 12 qualified days for the out-of-distribution (test-od) set. Input-Output Construction To ensure data quality and consistency across subgroups, we define a qualified patient-day as one with a complete 24-hour CGM trace (288 values) and adjacent continuity: the 24 hours prior and 8 hours following must also be complete. Each such day yields 24 hourly prediction input-output pairs, each comprising a 24-hour input and an 8-hour output segment. Each input-output pair must meet strict quality criteria: the input contains all 288 values, the output all 96; fewer than 40% of readings are constant (to exclude flatline readings); no values fall below 20 mg/dL; and fewer than 20% exceed 400 mg/dL. If any pair within a patient-day fails these checks, the entire day is discarded. This rigorous filtering ensures that every retained patient-day yields 24 high-quality input–output pairs, providing a reliable foundation for both classical and neural forecasting models. From each qualified day, we can extract 24 hourly prediction input-output pairs. Each pair consists of a 288-length input sequence (representing the prior 24 hours) and a 96-length output sequence (representing the subsequent 8 hours). The full output window supports evaluation up to 8 hours. The main text focuses on the first 2 hours of this window (24 steps), while results for 30-minute and 1-hour horizons are presented in Supplementary Note B. We also append static covariates (age group, gender, 12 diabetes type) and dynamic indicators (time-of-day) to each instance. Table 2 shows statistics: there are non-overlapping sets oftrain,valid,test-id, andtest-od, supporting evaluation across 12 balanced subgroups, with total 300 patients, and 5520 patient-days, and 132,480 input-output pairs. By enforcing uniform data quality and matched sample sizes, the partitioning scheme enables both robust generalization and fairness assessment. Detailed data descriptions, analysis, and accessibility information are provided in Supplementary Note A. Event Characterization FairGlucose includes timestamped behavioral event annotations logged via the mobile health platform: (1) meal events with nutritional details (carbs, calories, protein, fat, fiber), (2) exercise events with duration and intensity, and (3) medication events with dosage and timing. Of the 300 patients, 81 (27.0%) logged at least one event, producing 3,945 unique behavioral events: medication events (2,809; 71.2%), meal events (648; 16.4%), and exercise events (488; 12.4%). Logging patients recorded a median of 38 events (mean 48.7, range 1–237) over their observation period. Because the hourly sliding-window design produces overlapping 48-hour observation windows, each unique event appears in approximately 30 prediction pairs; consequently 22,239 of 132,480 samples (16.8%) contain at least one event within their window, yielding 117,979 event occurrences at the window level. Event timing is highly asymmetric: 33.8% of event occurrences fall in the 24-hour input window, 3.0% in the evaluated 2-hour forecast window, and 63.2% in the 2–24 hours after the prediction time. Because after-observation events are not necessarily known at deployment time, event-aware experiments quantify retrospective event-aware performance or an upper-bound estimate unless restricted to events available at prediction time. Demographic variations reveal age-related differences (patients 65+ log 22% fewer events than younger groups) and disease-type differences (T2D patients log 12% more events than T1D). This distinguishes FairGlucose from existing CGM benchmarks (which provide only glucose traces) and enables event-aware model development and subgroup-specific event impact analysis. Predictive Accuracy Root Mean Square Error (rMSE). rMSE measures the average magnitude of the prediction error across time steps and is sensitive to large deviations: rMSE = v u u t 1 T T X t=1 (ˆy t − y t ) 2 (1) Despite the right-skewed nature of rMSE, we report the mean to capture difficult cases and summarize overall model robustness [4]. Median values and other metrics are provided in Supplementary Note B for a fuller assessment. Clarke Error Grid Analysis (Clarke A): To assess clinical safety, we use the Clarke Error Grid [16], which categorizes predictions into five zones. Zone A represents the most clinically accurate predictions (within 20% or both<70 mg/dL) with no clinical errors, while Zone B captures benign errors with no treatment changes. Zones C, D, and E represent progressively dangerous errors: overcorrection (C), failure to detect hypo/hyperglycemia (D), and confusion between hypoglycemia and hyperglycemia (E). We report the Clarke A rate, representing predictions with perfect clinical accuracy. Critically, Zones D and E carry asymmetric clinical costs—hypoglycemia misses can lead to seizures or loss of consciousness within minutes, while hyperglycemia overcorrection may cause acute hypoglycemic events. This strict metric ensures models meet the highest clinical standards for glucose prediction validation [16, 26]. Statistical Uncertainty Because overlapping sliding windows produce multiple prediction pairs per patient, sample-level confidence intervals underestimate uncertainty. We therefore compute patient-level bootstrap confidence intervals: in each ofB = 1,000resamples, we draw patients with replacement, aggregate each patient’s per-sample rMSE to a patient-level mean, and compute the group statistic of interest (e.g., subgroup mean rMSE or between-group difference). The 95% CI is taken as the 2.5th and 97.5th percentiles of the bootstrap distribution. For subgroup comparisons (e.g., T1D vs. T2D, event 13 vs. no-event), we additionally report two-sided permutation tests (N = 10,000permutations) on the patient-level mean rMSE difference. All reportedp-values and confidence intervals in figures and tables use these patient-level procedures unless noted otherwise. Fairness Evaluation Frameworks To evaluate equity and robustness in glucose prediction models, we propose a fairness evaluation framework spanning five dimensions. Subgroup Fairness. Subgroup fairness examines whether predictive accuracy is consistent across patient demographics and clinical characteristics [27]. We partition the population intoKdisjoint subgroups (e.g., by age, sex, or diabetes type) and compute the average root mean squared error (rMSE i ) within each subgroupi. We quantify fairness disparities using two measures. First, the rMSE Spread captures the absolute difference between the best and worst performing subgroups. It is defined as rMSE-Spread = max i (rMSE i )− min i (rMSE i )(2) The second is the Gini Index [28] over the setrMSE 1 ,..., rMSE K , which measures the distribu- tional inequality in performance across subgroups. A higher Gini value indicates greater inequality in prediction error across subgroups, while a value closer to 0 indicates more equitable performance. Gini results are reported in Supplementary Note B. Input Length Sensitivity. Input length sensitivity assesses whether model performance improves as the input window length increases, and whether such gains are consistent across subgroups. Let L∈7, 37, 73, 145, 288denote the number of CGM input steps. For each lengthLand subgroupi, we define the temporal ratio (TR) as: TR (L) i = rMSE (L) i / rMSE (288) i (3) whererMSE (288) i is the performance under the longest (24-hour) input. A TR near 1.0 indicates consistent performance across input lengths; higher values indicate degradation with shorter input. Substantial variation in TR across subgroups at the same input length signals inequitable sensitivity to data availability; some subgroups may lose disproportionately more accuracy when historical data is limited. Instance-Level Difficulty. Instance-level difficulty investigates model robustness across prediction samples of varying difficulty. We define hard and easy instances using a vote-based approach across allMevaluated models [29]. For each model, we rank all samples by per-sample rMSE; samples in the top 25% (highest error) receive a hard vote, and samples in the bottom 25% (lowest error) receive an easy vote. We then compute vote countsVote hard j andVote easy j for each samplej. Samples with Vote hard j in the top 25th percentile of all vote counts are labeled hard; those withVote easy j in the top 25th percentile are labeled easy; the remainder are neutral. For each model, we compute rMSE over the hard and easy subsets separately. To assess robustness, we introduce normalized difficulty ratios using the Naive model as reference: Hard Ratio = rMSE hard model rMSE hard naive ,Easy Ratio = rMSE easy model rMSE easy naive (4) These ratios control for the inherent difficulty of the samples and emphasize how much a model improves upon a naive baseline in different difficulty regimes. A fair and robust model should yield reductions in both ratios, narrowing the performance gap between easy and hard cases. Event-Aware Evaluation. We compare models trained with event features (meals, exercise, medica- tion) versus without these features. Event features are evaluated under the availability caveat described above, because some events occur after prediction time. ComputingrMSE Event andrMSE NoEvent on the same test set, negative differences (∆ < 0) indicate event features improve accuracy. We evaluate across demographic subgroups to assess whether benefits are consistent and identify event-fairness disparities. Models To benchmark performance, we compare 33 models spanning four families. For statistical baselines (3 models), the Naive method predicts the most recent value, ARIMA [30] captures temporal 14 dependencies, and AutoETS fits exponential smoothing; all are implemented via Nixtla StatsForecast. For machine learning (4 models), we evaluate Linear Regression and tree-based models: LightGBM [18], CatBoost [31], and XGBoost [32], implemented via Nixtla MLForecast. For neural time-series models (20 architectures), we use the Time-Series-Library (TSLib) framework, which provides a unified training pipeline across diverse architectures. These include attention-based models: PatchTST [17], iTransformer, TimeXer, Crossformer, Transformer [33], Informer [34], Non-stationary Transformer, Autoformer [35], FEDformer, and TFT [36]; MLP/linear models: DLinear [37], TiDE, TSMixer, TimeMixer, N-BEATS, N-HiTS [38], and MICN; convolutional models: TimesNet and SCINet; and the zero-shot foundation model Chronos2. For frontier LLMs and foundation models (6 models), we include TimeGPT [19], a time-series-specific foundation model; GPT-5.1 [20] and GPT-5 Mini [21]; Claude 4.5 Sonnet and Claude 4.5 Haiku [22]; and Gemini 3 Flash [23]. These are evaluated via API-based inference. Experiment Details To ensure reproducibility, we trained and tested each model on predefined splits (train, validation, test-id, and test-od) with consistent input and output lengths. We focus on out-of-the-box performance to establish strong and fair baselines, deliberately avoiding model-specific enhancements such as pretraining, auxiliary losses, data augmentation, knowledge distillation, or learning rate schedules. Statistical and ML models use the Nixtla library [39] with Bayesian hyperparameter optimization via Optuna [40] (50 trials per model). Neural time-series models use the TSLib framework with default hyperparameters and up to 20-epoch training with early stopping (patience 5); see Supplementary Note C for search spaces. For LLMs, we set the temperature to zero for deterministic outputs with the same prompts. Experiment details are reported in Supplementary Note C. Our evaluation prioritizes general-purpose time-series forecasters over glucose-specific architectures [41,42] to assess out- of-the-box performance and facilitate broader adoption. Our framework establishes fairness-aware baselines that future work can extend with domain-specialized methods under the FairGlucose protocol. Ethics approval and consent to participate This research was performed in accordance with the Declaration of Helsinki. The secondary analysis of de-identified Welldoc patient data was reviewed and approved by the Johns Hopkins University Institutional Review Board (approval reference IRB00447704). No other ethics committee or institutional review board reviewed this work. Informed consent for the collection and research use of the underlying data was obtained from participants at the time of original data collection under Welldoc’s standard consent procedures. The Institutional Review Board waived the requirement for additional informed consent for the present secondary analysis because all records were fully de-identified before they were made available to the research team, so that participants could not be identified directly or through linked identifiers. Data availability The FairGlucose cohort metadata, demographic strata definitions, evaluation splits, and the 33-model reference leaderboard will be made available athttps://github.com/JHU-CDHAI/FairGlucose upon acceptance for publication. The underlying CGM and behavioral event records from the Welldoc cohort are not publicly available due to patient privacy; access can be requested under a data use agreement with Welldoc, Inc. Code availability All code for cohort construction, baseline training, evaluation, and figure generation will be made available athttps://github.com/JHU-CDHAI/FairGlucoseupon acceptance for publication. Pretrained model checkpoints used for the 33 baselines will be available on request from the corre- sponding author. 15 Acknowledgements The authors thank the Welldoc team for data access and clinical guidance. No specific funding was received for this work. Welldoc, Inc. provided access to the de-identified cohort and computing infrastructure but had no role in the study design, the analysis and interpretation of the results, the decision to publish, or the preparation of the manuscript. Author contributions J.L. and X.Z. led the cohort construction, data curation, software development, formal analysis, investigation, and visualization, and prepared the original draft of the manuscript. R.H. contributed to baseline-model implementation, data infrastructure, investigation, and manuscript review. A.K., A.K.I., and M.E.S. provided the Welldoc cohort and infrastructure access, contributed clinical and product expertise (validation), and reviewed and edited the manuscript. R.A. and G.G.G. contributed to conceptualization, supervision, and project administration, and reviewed and edited the manuscript. All authors reviewed and approved the final version of the manuscript. Competing interests R.H., A.K., A.K.I., and M.E.S. are employees of Welldoc, Inc. and may hold equity in the company. The remaining authors declare no competing interests. References [1]ElSayed, N. A. et al. 7. diabetes technology: standards of care in diabetes—2023. Diabetes Care 46, S111–S127 (2023). [2]Contreras, I. & Vehi, J. Artificial intelligence for diabetes management and decision support: literature review. Journal of medical Internet research 20, e10775 (2018). [3]Kim, H.-S. & Yoon, K.-H. Lessons from use of continuous glucose monitoring systems in digital healthcare. Endocrinology and Metabolism 35, 541–548 (2020). [4] Armandpour, M., Kidd, B., Du, Y. & Huang, J. Z. Deep personalized glucose level forecasting using attention-based recurrent neural networks. In 2021 International Joint Conference on Neural Networks (IJCNN), 1–8 (IEEE, 2021). [5]Luo, J. et al. A large sensor foundation model pretrained on continuous glucose monitor data for diabetes management. npj Health Systems 2, 35 (2025). [6]Mirshekarian, S., Shen, H., Bunescu, R. & Marling, C. Lstms and neural attention models for blood glucose prediction: Comparative experiments on real and synthetic data. In 2019 41st annual international conference of the IEEE engineering in medicine and biology society (EMBC), 706–712 (IEEE, 2019). [7] Subbaswamy, A. et al. A data-driven framework for identifying patient subgroups on which an AI/machine learning model may underperform. npj Digital Medicine 7 (2024). [8]Liu, M., Ning, Y., Teixayavong, S., Liu, X. et al. A scoping review and evidence gap analysis of clinical AI fairness. npj Digital Medicine 8 (2025). [9] Xu, Z., Yong, Y., Li, J., Yao, Q. & Zhou, S. K. Addressing fairness issues in deep learning-based medical image analysis: a systematic review. npj Digital Medicine 7 (2024). [10] Sergazinov, R. et al. Glucobench: Curated list of continuous glucose monitoring datasets with prediction benchmarks. arXiv preprint arXiv:2410.05780 (2024). [11] Marling, C. & Bunescu, R. The ohiot1dm dataset for blood glucose level prediction: Update 2020. In CEUR workshop proceedings, vol. 2675, 71 (2020). 16 [12]GluFormer Team. A generative foundation model for continuous glucose monitoring. Nature (2025). [13] Dubosson, F. et al. The open d1namo dataset: A multi-modal dataset for research on non- invasive type 1 diabetes management. Informatics in Medicine Unlocked 13, 92–100 (2018). [14]Hall, H. et al. Glucotypes reveal new patterns of glucose dysregulation. PLoS biology 16, e2005143 (2018). [15] Obermeyer, Z., Powers, B., Vogeli, C. & Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366, 447–453 (2019). [16]Clarke, W. L., Cox, D., Gonder-Frederick, L. A., Carter, W. & Pohl, S. L. Evaluating clinical accuracy of systems for self-monitoring of blood glucose. Diabetes care 10, 622–628 (1987). [17]Nie, Y., Nguyen, N. H., Sinthong, P. & Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. arXiv preprint arXiv:2211.14730 (2022). [18]Ke, G. et al. Lightgbm: A highly efficient gradient boosting decision tree. Advances in neural information processing systems 30 (2017). [19]Garza, A., Challu, C. & Mergenthaler-Canseco, M. Timegpt-1. arXiv preprint arXiv:2310.03589 (2023). [20]OpenAI. GPT-5.1 model (openai api documentation). OpenAI Platform Docs (2025). URL https://platform.openai.com/docs/models/gpt-5.1. Accessed 2026-01-29. [21]OpenAI. GPT-5 mini model (openai api documentation). OpenAI Platform Docs (2025). URL https://platform.openai.com/docs/models/gpt-5-mini. Accessed 2026-01-29. [22]Anthropic.What’s new in Claude 4.5 (claude api documentation).Claude API Docs (2025). URLhttps://platform.claude.com/docs/en/about-claude/models/ whats-new-claude-4-5. Accessed 2026-01-29. [23]Google. Gemini 3 developer guide (gemini api documentation). Google AI for Developers (2025). URLhttps://ai.google.dev/gemini-api/docs/gemini-3. Accessed 2026-01- 29. [24]Pikulin, S., Yehezkel, I. & Moskovitch, R. Enhanced blood glucose levels prediction with a smartwatch. Plos one 19, e0307136 (2024). [25]Sun, X., Rashid, M., Askari, M. R. & Cinar, A. Adaptive personalized prior-knowledge- informed model predictive control for type 1 diabetes. Control engineering practice 131, 105386 (2023). [26] Battelino, T. et al. Clinical targets for continuous glucose monitoring data interpretation: recommendations from the international consensus on time in range. Diabetes care 42, 1593– 1603 (2019). [27] Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K. & Galstyan, A. A survey on bias and fairness in machine learning. ACM computing surveys (CSUR) 54, 1–35 (2021). [28]Dixon, L., Li, J., Sorensen, J., Thain, N. & Vasserman, L. Measuring and mitigating unintended bias in text classification. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, 67–73 (2018). [29]Swayamdipta, S. et al. Dataset cartography: Mapping and diagnosing datasets with training dynamics. arXiv preprint arXiv:2009.10795 (2020). [30]Yang, J., Li, L., Shi, Y. & Xie, X. An arima model with adaptive orders for predicting blood glucose concentrations and hypoglycemia. IEEE journal of biomedical and health informatics 23, 1251–1260 (2018). 17 [31]Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V. & Gulin, A. Catboost: unbiased boosting with categorical features. Advances in neural information processing systems 31 (2018). [32]Chen, T. & Guestrin, C. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, 785–794 (2016). [33]Vaswani, A. et al. Attention is all you need. Advances in neural information processing systems 30 (2017). [34]Zhou, H. et al. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, vol. 35, 11106–11115 (2021). [35]Wu, H., Xu, J., Wang, J. & Long, M. Autoformer: Decomposition transformers with auto- correlation for long-term series forecasting. Advances in neural information processing systems 34, 22419–22430 (2021). [36]Lim, B., Arık, S. Ö., Loeff, N. & Pfister, T. Temporal fusion transformers for interpretable multi-horizon time series forecasting. International journal of forecasting 37, 1748–1764 (2021). [37] Zeng, A., Chen, M., Zhang, L. & Xu, Q. Are transformers effective for time series forecasting? In Proceedings of the AAAI conference on artificial intelligence, vol. 37, 11121–11128 (2023). [38]Challu, C. et al. Nhits: Neural hierarchical interpolation for time series forecasting. In Proceedings of the AAAI conference on artificial intelligence, vol. 37, 6989–6997 (2023). [39] Olivares, K. G., Challú, C., Garza, A., Canseco, M. M. & Dubrawski, A. NeuralForecast: User friendly state-of-the-art neural forecasting models. PyCon Salt Lake City, Utah, US 2022 (2022). URL https://github.com/Nixtla/neuralforecast. [40]Akiba, T., Sano, S., Yanase, T., Ohta, T. & Koyama, M. Optuna: A next-generation hyper- parameter optimization framework. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, 2623–2631 (2019). [41]Sergazinov, R., Armandpour, M. & Gaynanova, I. Gluformer: Transformer-based personalized glucose forecasting with uncertainty quantification. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5 (IEEE, 2023). [42]Zhu, T., Li, K., Herrero, P. & Georgiou, P. Personalized blood glucose prediction for type 1 diabetes using evidential deep learning and meta-learning. IEEE Transactions on Biomedical Engineering 70, 193–204 (2022). [43]Pedregosa, F. et al. Scikit-learn: Machine learning in python. Journal of Machine Learning Research 12, 2825–2830 (2011). [44] Garza, A., Canseco, M. M., Challú, C. & Olivares, K. G. StatsForecast: Lightning fast forecasting with statistical and econometric models. PyCon Salt Lake City, Utah, US 2022 (2022). URL https://github.com/Nixtla/statsforecast. [45]Oreshkin, B. N., Carpov, D., Chapados, N. & Bengio, Y. N-beats: Neural basis expansion analysis for interpretable time series forecasting. In International Conference on Learning Representations (2020). 18 Dataset Statistics and Details FairGlucose Dataset Statistics For each subgroup in the dataset (e.g., by age, gender, diabetes type), we computed key signal metrics such as mean, median, standard deviation, minimum, maximum, coefficient of variation (CV), interquartile range (IQR), and selected percentiles (5th, 25th, 75th, 95th). Additionally, we calculated clinically relevant measures including TIR% (Time in Range), MAGE (Mean Amplitude of Glycemic Excursions), and dynamic features such as the average rate of change, number of peaks/valleys, average trend length, and entropy of the signal distribution. The implementation usedscipy.stats.entropyfor entropy computation andscipy.signal.find_peaksto detect peaks and valleys [43]. The results were aggregated to provide an overview of the variability and distribution of glucose values across subsets. For all calculations, each metric was first computed at the individual sequence level using a window of 288 CGM (corresponding to a full-day CGM trace). We then aggregated these results by taking the mean across all sequences within the subgroup to obtain the representative metric value for that subset. Formally, the metrics are defined as follows: • Interquartile Range (IQR): IQR = Q 75 − Q 25 (5) where Q p denotes the p-th percentile. • Selected Percentiles: Values at 5th, 25th, 75th, and 95th percentiles of the sequence. • Time in Range (TIR%): TIR% = #x i | L≤ x i ≤ U n × 100(6) whereLandUare clinically defined lower and upper bounds (typically 70 and 180 mg/dL). •Mean Amplitude of Glycemic Excursions (MAGE): The average of all significant excur- sions exceeding one standard deviation from the mean: MAGE = P m k=1 |e k | m , e k > σ(7) where e k are individual excursions and m is the number of qualifying excursions. • Average Rate of Change: Rate of Change = 1 n− 1 n−1 X i=1 |x i+1 − x i | ∆t (8) where ∆t is the time interval between measurements. •Number of Peaks / Valleys: The count of local maxima and minima in the sequence, detected using scipy.signal.find_peaks. • Average Trend Length: The mean duration of monotonic segments between peaks and valleys. • Entropy: Shannon entropy of the empirical distribution: H =− k X j=1 p j log(p j )(9) wherep j is the probability of values falling in binj,computed using scipy.stats.entropy. Cohort Representation in Feature Space To verify that the 12 demographic strata constitute distinguishable subpopulations in the input space (rather than balanced groups that nevertheless overlap in CGM dynamics), we visualized 2-D t-SNE embeddings of 288-step 24-hour patient-day CGM traces under five subgroup-coloring schemes 19 806040200204060 t-SNE Component 1 40 20 0 20 40 t-SNE Component 2 t-SNE Visualization of input_ids by stratum 40-64_1_1.0 18-39_1_1.0 65+_2_2.0 40-64_2_1.0 18-39_2_1.0 65+_1_1.0 65+_2_1.0 40-64_1_2.0 18-39_1_2.0 40-64_2_2.0 65+_1_2.0 18-39_2_2.0 (a) By stratum (age× gender× type) 806040200204060 t-SNE Component 1 40 20 0 20 40 t-SNE Component 2 t-SNE Visualization of input_ids by age 40-64 18-39 65+ (b) By age 806040200204060 t-SNE Component 1 40 20 0 20 40 t-SNE Component 2 t-SNE Visualization of input_ids by gender 1 2 (c) By gender 806040200204060 t-SNE Component 1 40 20 0 20 40 t-SNE Component 2 t-SNE Visualization of input_ids by type 1.0 2.0 (d) By diabetes type Figure 7: t-SNE embeddings of 24-hour CGM patient-day traces (288 5-minute steps), colored by (top) composite stratum and (bottom) univariate attributes. Stratum-coloring produces the most separable clusters; univariate colorings overlap substantially. Star markers denote subgroup centroids. (Figure 7). The composite stratum-coloring (age×gender×diabetes type) produces the most pronounced cluster separation; univariate colorings by age, gender, or diabetes type alone show substantially overlapping clusters. This confirms that the cohort’s multi-dimensional stratification captures structure that univariate stratification would miss, and supports the interpretation that down- stream model-class-invariant disparities reflect genuine inter-subgroup differences in the prediction task rather than nominal labeling. Visualizing Feature Correlations and Subgroup Differences in CGM Data To investigate the distributional characteristics of CGM sequence features across different patient subgroups, we computed a set of descriptive statistics for each 24-hour sequence. These include: mean glucose, standard deviation (std), slope (reflecting directional change), number of peaks (n_peaks), and clinically meaningful measures such as Time-In-Range (TIR), Time-Below-Range (TBR), and Time-Above-Range (TAR). These features were extracted using a custom feature engineering function (advanced_sequence_features) applied uniformly across all sequences. We visualized the relationships among these features usingseaborn.pairplotwithcorner=True, stratified by various patient attributes such as stratum, age group, gender, diabetes type, and dataset split. Diagonal plots show kernel density estimates (KDEs) for feature distributions, while off- diagonal plots display pairwise scatter plots, allowing for the examination of feature correlations and cluster separability across groups. Our exploratory analysis reveals that certain features are strongly correlated. Notably, mean glucose exhibits a strong positive correlation with TAR and an inverse relationship with TIR, reflecting the direct impact of average glucose levels on clinical range metrics. The number of peaks and slope—features that capture temporal dynamics—exhibit weaker correlations with mean or range- based metrics but still capture unique behavioral signatures of glucose traces. 20 Importantly, we observe substantial differences in the distributional shapes and feature interactions across strata. For example, younger age groups (18–39) tend to show wider dispersion in both mean glucose and TIR, indicating greater physiological variability or lifestyle-driven fluctuations. In contrast, some strata, such as older patients with T2D, demonstrate narrower distributions, possibly reflecting more stable or medically regulated glucose profiles. These cross-stratum differences highlight the importance of subgroup-sensitive modeling. From a modeling perspective, these patterns have direct implications. High intra-stratum variability can challenge prediction consistency, while inter-stratum distributional shifts may introduce bias if not properly addressed. Such differences motivate the design of fairness-aware models that explicitly account for heterogeneous feature landscapes. Moreover, the scatter plots also reveal outlier patterns and nonlinear dependencies not easily captured by global metrics. For instance, in some strata, the relationship between TIR and n_peaks is non-monotonic, suggesting that overly simplistic feature assumptions may fail in real-world deployment. Resource Availability and Use FairGlucose is provided through two complementary access mechanisms designed to balance research utility with patient-data confidentiality. Method 1: Controlled access via Data Use Agreement (DUA). Researchers may request access to the de-identified CGM cohort by submitting a written research proposal, institutional affiliation, planned analyses, and IRB approval or exemption. Approved requests proceed to a formal DUA between the requesting institution and Welldoc, Inc. that governs scope of use, data security, publi- cation policy, and redistribution restrictions. Access is provisioned within a secure computational environment that enables interactive analysis while preventing unauthorized export of patient-level data. Method 2: API-based evaluation server. For investigators who require benchmarking without raw-data access, we provide an API endpoint that accepts trained forecasting models and returns standardized metrics (rMSE, MAE, Clarke Error Grid zone counts, and subgroup-disaggregated variants) on the held-out test set. Each submitted model runs in a sandboxed environment and is never exposed to user-controlled execution paths; results are returned in a uniform machine-readable report. This model-to-data pattern preserves data privacy while enabling reproducible third-party comparison against the reference 33-model leaderboard. Cohort metadata, demographic strata definitions, evaluation splits, and the 33-model reference leaderboard will be made openly available athttps://github.com/JHU-CDHAI/FairGlucose upon acceptance. Both access mechanisms ensure that FairGlucose can support the broadest possible community of investigators while protecting the patients whose data made the benchmark possible. Model Evaluation Details Expanded Model Evaluation Across Forecast Horizons This section presents accuracy comparisons across different forecast horizons (30 minutes, 1 hour, and 2 hours) using three metrics: MAE, rMSE, and TIR Absolute Error [4, 26]. To complement the main results, we present detailed performance metrics for all 33 models across three prediction horizons, 30 minutes, 1 hour, and 2 hours, evaluated on both in-distribution (Test-ID) and out-of-distribution (Test-OD) sets. Each horizon-specific evaluation reports root mean squared error (rMSE), mean absolute error (MAE), and Time-in-Range Absolute Error (TIR-AE), capturing both numerical and clinically relevant dimensions of model performance. The corresponding results are reported in the tables below. At the 30-minute horizon, attention-based neural models consistently achieve the strongest perfor- mance across all three metrics and test splits. NS-Transformer and PatchTST [17] rank among the top performers, with rMSE and MAE among the lowest of all models. N-HiTS performs similarly well, especially in the MAE metric. TFT [36] also delivers competitive results at this horizon, though with slightly higher TIR error. Among classical models, ARIMA [30] shows surprisingly strong performance, often outperforming machine learning methods like XGBoost [32] and CatBoost [31] in TIR error, particularly on the Test-OD set. LightGBM [18] and CatBoost, while strong in rMSE 21 and MAE, show more variability in TIR prediction across subgroups. Foundation models, including GPT-5.1 [20], Gemini 3 Flash [23], and TimeGPT [19], consistently underperform across all 30- minute evaluations. Their rMSE and TIR error are notably higher than neural and classical models, often close to or worse than the naive baseline. These results indicate that large-scale pretraining alone is insufficient for accurate short-term physiological forecasting. Moving to the 1-hour horizon, the performance gap between top and bottom models becomes more pronounced. Attention-based architectures (NS-Transformer, PatchTST, TimeXer) maintain their lead in both error-based and TIR-based metrics. N-HiTS and TFT continue to perform competitively, though their relative rankings fluctuate slightly depending on the metric and test split. As the prediction window increases, statistical and machine learning models begin to fall further behind. ARIMA’s relative strength diminishes, particularly on the Test-OD set, where its TIR error increases more sharply than the neural baselines. The degradation in performance is more severe for foundation models, which show clear limitations in both MAE and TIR prediction. These results highlight the challenges of mid-horizon glucose prediction and reaffirm the superiority of specialized neural time-series models for this task. At the 2-hour horizon, performance deteriorates further across all models, as expected given the increased forecasting difficulty. Despite this, NS-Transformer maintains the best overall rMSE (25.6 mg/dL), followed closely by Transformer (25.7), TimeXer and PatchTST (25.8). TFT remains strong but exhibits slightly higher TIR error, suggesting increased difficulty in capturing clinical relevance at extended horizons. In contrast, all classical models, including ARIMA, linear regression, and tree-based methods, exhibit larger declines in performance, particularly in TIR error, indicating limited capacity to generalize over long-term physiological dynamics. Foundation models continue to underperform substantially. Their rMSE and TIR metrics are consistently worse than even the simplest baselines, underscoring their inadequacy for long-horizon CGM prediction without targeted adaptation or fine-tuning. The widening performance gap reinforces the importance of task-specific architecture and domain-aware training for reliable long-range forecasting. Across all horizons, comparison between Test-ID and Test-OD performance reveals remarkable stability. For nearly all models, the performance degradation under distribution shift is minimal, with OD/ID error ratios typically close to 1.0. This consistency suggests that our dataset construction and evaluation framework provide a robust test bed for assessing real-world generalization. Top- performing models, particularly NS-Transformer, PatchTST, and TFT, demonstrate strong resilience to distributional variation, maintaining their relative rankings across both splits. In contrast, models with high Test-ID errors tend to generalize poorly to the Test-OD set as well, reinforcing the observation that pretraining alone does not confer robustness in this domain. Overall, the extended results affirm the main findings of the benchmark. Attention-based neural architectures (NS-Transformer, PatchTST, TimeXer) are the most accurate and robust models across horizons and evaluation settings. N-HiTS and TFT also perform strongly, especially under longer horizons and distribution shifts. Classical models offer solid baselines, particularly at short horizons, while foundation models fail to generalize effectively without adaptation. These observations emphasize the value of rigorous benchmark design and highlight the importance of domain-specific modeling approaches for fair and reliable glucose prediction. Expanded results from Figures 12, 13, and 14 also present the similar pattern. Moreover, for each time horizon, the figures illustrate the performance variation among different data splits (Test-ID and Test-OD). Shorter horizons (30 minutes) exhibit lower errors across all metrics, while errors increase as the prediction window extends to 1 hour and 2 hours, reflecting the expected decline in forecast accuracy with longer lead times. These visualizations highlight consistent patterns across metrics. Fairness Evaluation: Subgroup Heterogeneity Figure 15 presents a comprehensive comparison of model performance across demographic and clinical subgroups for all three prediction horizons (30 minutes, 1 hour, and 2 hours). Each subfigure displays rMSE performance across seven subgroups organized by dimension: three age groups (18–39, 40–64, 65+), two gender categories (male, female), and two diabetes types (T1D, T2D). Six representative models are shown spanning statistical (ARIMA), machine learning (LightGBM), neural time-series (PatchTST, TFT), and frontier LLM (GPT-5.1, TimeGPT) approaches. 22 The visualization employs a two-panel design to assess both accuracy and generalization. The top panel shows absolute Test-ID rMSE performance with consistent y-axis ranges across horizons (30m: 5–15 mg/dL, 1h: 10–25 mg/dL, 2h: 20–35 mg/dL), enabling direct cross-horizon comparison of error magnitudes. Grouped bar charts allow within-subgroup model comparison, with vertical separators delineating the three demographic dimensions. The bottom panel uses scatter plots to visualize OD/ID error ratios, where values near 1.0 indicate robust generalization and deviations reveal subgroup-specific distribution shift vulnerabilities. A green shaded band (0.95–1.05) highlights the region of near-perfect generalization. Consistent with the patterns discussed in the main text, several intersectional disparities emerge. Male patients generally exhibit better accuracy than female patients in the 18–39 age group, while the reverse is observed in the 65+ group, indicating age- and gender-specific interactions in prediction difficulty. T1D patients consistently show higher rMSE than T2D patients across horizons, reaffirming that T1D populations are more challenging to forecast despite their structured insulin regimens. Neural time-series models (NS-Transformer, PatchTST, TFT) demonstrate more consistent performance across subgroups compared to frontier LLMs, which exhibit greater variability, particularly for T1D and younger age groups. The OD/ID ratios in the lower panels further expose subgroup-specific vulnerabilities. While the overall generalization performance appears stable in population-level evaluations (e.g., Figures 12– 14), subgroup-level ratios span a broader range, consistent with the 0.8–1.4 range reported in the main text. Notably, most models maintain ratios near 1.0 for T2D patients across all age groups, indicating robust generalization for this population. In contrast, T1D patients, particularly females aged 40–64, show more pronounced distribution shift effects. Frontier LLMs exhibit more variable generalization patterns depending on subgroup, suggesting limited robustness to demographic shifts. As prediction horizons extend from 30 minutes to 2 hours, absolute errors increase predictably across all subgroups (note the increasing y-axis ranges), but relative subgroup disparities remain largely stable. This consistency suggests that subgroup-specific challenges are inherent to the prediction task rather than artifacts of model capacity or training. However, the generalization patterns (bottom panels) show slight deterioration at longer horizons, with OD/ID ratio variance increasing from 0.05 at 30 minutes to 0.08 at 2 hours, indicating that distribution shift effects compound with forecast difficulty. Together, these appendix results substantiate the main text’s findings: generalization robustness and predictive accuracy vary significantly across demographic and clinical subgroups. Fairness-aware evaluations should therefore extend beyond marginal attributes to capture intersectional effects, ensuring models perform reliably across all subpopulations in real-world deployments. The consistent y-axis scaling and scatter plot visualization facilitate rapid identification of vulnerable subgroups and poorly generalizing models, supporting targeted model improvement and fairness-aware deployment decisions. Fairness Evaluation: Model Accuracy vs Fairness To evaluate subgroup fairness across models, we computed two measures of variability: Prediction Error Spread (the difference between maximum and minimum subgroup errors, shown in Figure 16) and Gini coefficient (shown in Figure 17). For each model, we first calculated the root mean squared error (rMSE) across all subgroups, followed by the fairness metric that captures how evenly errors are distributed across these subgroups. A lower Spread or Gini metric indicates greater fairness (less disparity). We generated separate plots for 30-minute, 1-hour, and 2-hour horizons, using both spread and Gini fairness measures. Each point in the scatter plots represents a model, with the x-axis showing average predictive error (lower is better) and the y-axis representing fairness disparity (lower is fairer). Models were grouped by model family. The fairness analysis reveals notable differences among model families. Neural time-series models such as PatchTST and TFT occupy favorable regions with low error and low subgroup disparity, consistent with the main-text accuracy–fairness analysis. Conversely, several lower-performing neural and frontier LLM models exhibit higher error and/or higher subgroup disparity, indicating less consistent performance across subgroups. Traditional linear 23 models and statistical baselines occupy the middle ground with moderate fairness and error. Overall, the tradeoff analysis supports reporting both average accuracy and subgroup disparity rather than selecting models by population-level rMSE alone. Input Length Sensitivity: Extended Results The main text presents input-length sensitivity analysis for the 2-hour prediction horizon. Here we provide extended results across all three horizons (30 minutes, 1 hour, and 2 hours) using three metrics: rMSE, MAE, and TIR Absolute Error. Figure 8, panels (a)–(f), covers rMSE and MAE at all three horizons, and Figure 9, panels (a)–(c), covers TIR Absolute Error. Together they show how model performance varies across input lengths for each horizon-metric combination. Across all settings, the same subgroup-dependent pattern holds: younger T1D patients are relatively insensitive to input-length reduction, while older T2D patients show larger degradation with shorter histories. Neural models (PatchTST, N-HiTS, TFT) consistently benefit most from longer input windows, while statistical and foundation models show more modest gains. 73773145288 Input Length (steps) 10 15 20 25 30 35 30m rMSE (mg/dL) Input Length Sensitivity: 30m rMSE Naive ARIMA AutoETS NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer N-BEATS FEDformer Autoformer Chronos2 (a) 30-minute rMSE 73773145288 Input Length (steps) 20 25 30 35 1h rMSE (mg/dL) Input Length Sensitivity: 1h rMSE Naive ARIMA AutoETS NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer N-BEATS FEDformer Autoformer Chronos2 (b) 1-hour rMSE 73773145288 Input Length (steps) 26 28 30 32 34 36 38 40 2h rMSE (mg/dL) Input Length Sensitivity: 2h rMSE Naive ARIMA AutoETS NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer N-BEATS FEDformer Autoformer Chronos2 (c) 2-hour rMSE 73773145288 Input Length (steps) 10 15 20 25 30 35 30m MAE (mg/dL) Input Length Sensitivity: 30m MAE Naive ARIMA AutoETS NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer N-BEATS FEDformer Autoformer Chronos2 (d) 30-minute MAE 73773145288 Input Length (steps) 15 20 25 30 35 1h MAE (mg/dL) Input Length Sensitivity: 1h MAE Naive ARIMA AutoETS NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer N-BEATS FEDformer Autoformer Chronos2 (e) 1-hour MAE 73773145288 Input Length (steps) 22 24 26 28 30 32 34 36 38 2h MAE (mg/dL) Input Length Sensitivity: 2h MAE Naive ARIMA AutoETS NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer N-BEATS FEDformer Autoformer Chronos2 (f) 2-hour MAE Figure 8: Input length sensitivity across prediction horizons and metrics. Each panel shows model performance (y-axis) at varying input lengths (x-axis) for a specific horizon-metric combination. 73773145288 Input Length (steps) 5.0 7.5 10.0 12.5 15.0 17.5 20.0 22.5 25.0 30m TIR_error (%) Input Length Sensitivity: 30m TIR_error Naive ARIMA AutoETS NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer N-BEATS FEDformer Autoformer Chronos2 (a) 30-minute TIR-AE 73773145288 Input Length (steps) 8 10 12 14 16 18 20 22 24 1h TIR_error (%) Input Length Sensitivity: 1h TIR_error Naive ARIMA AutoETS NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer N-BEATS FEDformer Autoformer Chronos2 (b) 1-hour TIR-AE 73773145288 Input Length (steps) 14 16 18 20 22 24 26 2h TIR_error (%) Input Length Sensitivity: 2h TIR_error Naive ARIMA AutoETS NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer N-BEATS FEDformer Autoformer Chronos2 (c) 2-hour TIR-AE Figure 9: Input length sensitivity for TIR Absolute Error across prediction horizons. 24 Instance-Level Difficulty: Extended Results The main text presents instance-level difficulty analysis for the 2-hour horizon. Here we provide extended difficulty-stratified results across all three horizons. Figure 10, panels (a)–(f), covers rMSE and MAE at all three horizons, and Figure 11, panels (a)–(c), covers TIR Absolute Error. Together they show the subgroup-level distribution of hard and easy cases and the corresponding rMSE spread across demographic strata for each horizon-metric combination. The key finding from the main text—that subgroup disparities align with the proportion of hard cases—holds consistently across all horizons. Young male T1D and middle-aged female T1D groups maintain the highest hard-case proportions regardless of the prediction window. M T1D F T1D M T2D F T2D 0 10 20 30 40 50 60 70 80 Mean rMSE Age: 18-39 easy medium hard M T1D F T1D M T2D F T2D Age: 40-64 M T1D F T1D M T2D F T2D Age: 65+ Hard/Easy Case rMSE Spread: 30m (a) 30-minute rMSE M T1D F T1D M T2D F T2D 0 10 20 30 40 50 60 70 80 Mean rMSE Age: 18-39 easy medium hard M T1D F T1D M T2D F T2D Age: 40-64 M T1D F T1D M T2D F T2D Age: 65+ Hard/Easy Case rMSE Spread: 1h (b) 1-hour rMSE M T1D F T1D M T2D F T2D 0 10 20 30 40 50 60 70 80 Mean rMSE Age: 18-39 easy medium hard M T1D F T1D M T2D F T2D Age: 40-64 M T1D F T1D M T2D F T2D Age: 65+ Hard/Easy Case rMSE Spread: 2h (c) 2-hour rMSE M T1D F T1D M T2D F T2D 0 10 20 30 40 50 60 70 80 Mean MAE Age: 18-39 easy medium hard M T1D F T1D M T2D F T2D Age: 40-64 M T1D F T1D M T2D F T2D Age: 65+ Hard/Easy Case MAE Spread: 30m (d) 30-minute MAE M T1D F T1D M T2D F T2D 0 10 20 30 40 50 60 70 80 Mean MAE Age: 18-39 easy medium hard M T1D F T1D M T2D F T2D Age: 40-64 M T1D F T1D M T2D F T2D Age: 65+ Hard/Easy Case MAE Spread: 1h (e) 1-hour MAE M T1D F T1D M T2D F T2D 0 10 20 30 40 50 60 70 80 Mean MAE Age: 18-39 easy medium hard M T1D F T1D M T2D F T2D Age: 40-64 M T1D F T1D M T2D F T2D Age: 65+ Hard/Easy Case MAE Spread: 2h (f) 2-hour MAE Figure 10: Instance-level difficulty analysis: subgroup-level hard/easy case distribution and rMSE/- MAE spread across prediction horizons. M T1D F T1D M T2D F T2D 0 10 20 30 40 50 60 70 80 Mean TIR_error Age: 18-39 easy medium hard M T1D F T1D M T2D F T2D Age: 40-64 M T1D F T1D M T2D F T2D Age: 65+ Hard/Easy Case TIR_error Spread: 30m (a) 30-minute TIR-AE M T1D F T1D M T2D F T2D 0 10 20 30 40 50 60 70 80 Mean TIR_error Age: 18-39 easy medium hard M T1D F T1D M T2D F T2D Age: 40-64 M T1D F T1D M T2D F T2D Age: 65+ Hard/Easy Case TIR_error Spread: 1h (b) 1-hour TIR-AE M T1D F T1D M T2D F T2D 0 10 20 30 40 50 60 70 80 Mean TIR_error Age: 18-39 easy medium hard M T1D F T1D M T2D F T2D Age: 40-64 M T1D F T1D M T2D F T2D Age: 65+ Hard/Easy Case TIR_error Spread: 2h (c) 2-hour TIR-AE Figure 11: Instance-level difficulty analysis for TIR Absolute Error across prediction horizons. Subgroup-Level Performance Tables Tables 5–8 present the full subgroup-disaggregated performance for all 33 models at the 2-hour prediction horizon, evaluated on both in-distribution (Test-ID) and out-of-distribution (Test-OD) splits. Models are grouped by family (Statistical, ML, Neural, API) and sorted by overall rMSE within each family. These tables complement the main results (Table 1 of the main text) by providing per-subgroup breakdowns by diabetes type (T1D, T2D), gender (Female, Male), and age group (18–39, 40–64, 65+). 25 Table 5: 2-hour rMSE (mg/dL) across subgroups – Test-ID. Lower values indicate better prediction accuracy. T1D patients consistently show higher rMSE than T2D across all models, reflecting greater glucose variability. ModelAllT1DT2DFemaleMale18-3940-6465+ Statistical ARIMA27.3930.8423.9427.3627.4228.4826.9326.70 Naive29.0332.3525.7028.9629.1029.9628.8428.19 AutoETS29.8133.6126.0129.8729.7530.7929.8128.69 ML Linear25.8729.1522.6025.7925.9627.1325.7024.65 LightGBM25.8929.3822.4025.6526.1327.2425.7324.55 XGBoost27.6631.1024.2227.8427.4729.0227.5326.27 CatBoost28.4731.5125.4228.4828.4629.4328.2027.69 Neural PatchTST25.5428.7522.3325.4125.6726.4825.5624.44 NS-Trans.25.5628.6322.4925.4925.6326.6825.5024.35 Transformer25.5728.7022.4425.4725.6726.7325.4424.40 TimeXer25.6128.7522.4625.4925.7326.6925.4924.52 Crossformer25.6628.7522.5625.6625.6526.6425.5424.68 Informer25.8028.9822.6225.6725.9226.9425.5824.76 TFT25.8228.9122.7425.6825.9726.7125.6325.04 MICN25.9729.1022.8425.7626.1827.0625.9124.81 DLinear26.2129.3923.0325.9526.4827.4025.9525.17 TimesNet26.2329.6922.7826.1826.2927.5826.0224.95 SCINet26.3029.4323.1726.0926.5127.5526.0025.24 iTransformer26.5029.8523.1526.1926.8127.5526.5725.21 TSMixer26.5829.8523.3126.3826.7827.5426.6125.46 TiDE26.6029.8123.4026.2926.9127.7226.3225.66 TimeMixer26.6129.8023.4226.3526.8827.7426.4525.53 N-HiTS26.6530.1623.1326.2127.0827.5726.6425.61 N-BEATS27.2430.6723.8126.9627.5228.2526.9026.51 FEDformer30.5334.4926.5830.3130.7631.4830.0630.02 Autoformer34.8539.5730.1334.6635.0435.5233.7435.41 Chronos239.0544.1733.9438.5839.5240.1237.6339.53 API TimeGPT26.5429.8323.2526.5226.5727.5226.3025.72 GPT-5.127.2730.4624.0827.2127.3328.6727.3625.58 GPT-5 Mini27.5730.7424.4127.5727.5828.5027.6426.43 Claude 4.527.8930.6725.1128.1927.6028.5028.6626.29 Claude 4.5H28.4731.4825.4628.5028.4429.6228.5427.08 Gemini 330.7033.5027.8930.6630.7432.1230.5329.27 26 Table 6: 2-hour rMSE (mg/dL) across subgroups – Test-OD. Out-of-distribution evaluation on held- out patients. OD/ID ratios near 1.0 indicate robust generalization; subgroup-level variation reveals hidden heterogeneity. ModelAllT1DT2DFemaleMale18-3940-6465+ Statistical ARIMA27.5030.7324.2628.4326.5727.0429.1626.33 Naive29.0232.4925.5530.1127.9328.8630.7327.51 AutoETS29.6333.0826.1831.0028.2629.7631.2427.92 ML LightGBM26.0228.7623.2927.0225.0325.9427.4424.72 Linear26.2028.7023.7027.2425.1726.4727.3124.85 XGBoost27.9630.8625.0628.9826.9327.1229.9426.86 CatBoost28.4631.6125.3129.6427.2828.7729.9326.69 Neural NS-Trans.25.6928.2923.0826.5624.8225.2927.2824.53 Transformer25.7428.4323.0526.5824.8925.4727.4824.29 TimeXer25.8928.3123.4626.7625.0225.2527.4525.00 PatchTST25.9728.6023.3426.7525.1925.4927.5724.89 Crossformer26.0428.6523.4426.9225.1625.4727.6625.03 Informer26.0828.6023.5626.9125.2525.3027.7225.27 TFT26.2128.6623.7627.1125.3025.7127.8625.09 MICN26.2528.8223.6827.1425.3525.6727.8525.26 iTransformer26.4929.1423.8427.2825.7026.0227.8125.67 TimesNet26.5729.2123.9227.3225.8125.7728.0125.97 DLinear26.5929.1824.0027.3525.8425.8728.1125.85 SCINet26.7029.2124.1827.5625.8426.0628.1925.89 N-HiTS26.9129.9723.8427.7026.1126.6127.9726.17 TiDE26.9929.6124.3727.8226.1726.1328.5226.37 TSMixer27.0129.8524.1827.7026.3326.3328.4626.29 TimeMixer27.2530.0824.4228.0826.4226.6828.8426.26 N-BEATS27.6130.4624.7628.3526.8727.0529.1026.72 FEDformer31.8134.8528.7632.8630.7530.5633.2931.62 Autoformer36.4239.8532.9938.0234.8234.1938.3136.82 Chronos240.3444.2636.4241.7038.9837.5941.7841.72 API TimeGPT26.7229.6523.8028.0625.3926.4128.4725.32 GPT-5.127.8630.7724.9528.8526.8727.1230.3026.23 GPT-5 Mini–26.6127.31–26.02 Claude 4.528.5131.0725.9529.5427.4828.5030.7726.31 Claude 4.5H29.0732.0826.0630.1228.0328.6831.1427.44 Gemini 331.8834.8028.9533.1130.6430.4234.8430.44 27 Table 7: 2-hour MAE (mg/dL) across subgroups – Test-ID. Mean absolute error provides a comple- mentary view to rMSE, being less sensitive to large outlier predictions. ModelAllT1DT2DFemaleMale18-3940-6465+ Statistical ARIMA23.3526.2220.4923.3123.3924.2722.9522.78 Naive24.7827.5522.0124.6724.9025.5424.6124.13 AutoETS25.2428.3522.1225.2425.2326.0525.2624.28 ML Linear22.0324.7719.2921.9922.0723.1221.8520.99 LightGBM21.9724.9019.0321.8422.1023.1521.8320.79 XGBoost23.4026.2820.5223.5723.2324.6023.2722.18 CatBoost24.2626.7821.7424.2624.2625.0224.0323.66 Neural PatchTST21.8724.5819.1721.8121.9322.7021.8320.97 NS-Trans.21.7824.3419.2321.7821.7922.7921.7020.75 Transformer21.8324.4219.2321.7821.8722.8521.6820.84 TimeXer21.9524.5919.3121.9022.0022.8821.7921.07 Crossformer21.8824.4519.3121.9321.8322.7421.7421.08 Informer21.9924.6619.3221.9222.0523.0221.7421.11 TFT21.9724.5019.4421.9122.0422.7321.7821.34 MICN22.1324.7519.5022.0022.2623.1121.9921.18 DLinear22.3825.0519.7122.1822.5823.4422.0721.54 TimesNet22.5325.5219.5522.5222.5523.7322.2421.53 SCINet22.4125.0119.8022.2722.5423.5322.0621.54 iTransformer22.8125.6319.9822.5623.0523.7622.8521.68 TSMixer22.9525.7020.2022.8423.0723.8222.9022.03 TiDE22.8425.5420.1522.6123.0723.8222.5222.12 TimeMixer22.9925.7020.2822.8023.1823.9822.8022.09 N-HiTS22.9825.9719.9822.6523.3023.8222.8822.13 N-BEATS23.7226.6820.7723.5023.9424.5223.3923.22 FEDformer27.1930.7523.6227.0427.3427.9826.7426.82 Autoformer31.3035.6826.9231.1431.4631.8130.1632.07 Chronos236.0440.7931.2935.5536.5336.9834.5536.73 API TimeGPT22.5125.2619.7622.4922.5223.3622.2921.81 GPT-5.123.0525.7220.3923.0423.0724.2923.1021.58 GPT-5 Mini23.2225.8220.6223.2223.2223.9723.3222.24 Claude 4.523.4825.7521.2023.7823.1823.9724.2022.06 Claude 4.5H23.9626.4321.4924.0123.9124.9424.0222.77 Gemini 325.9428.2623.6325.9525.9327.2125.8324.64 28 Table 8: 2-hour MAE (mg/dL) across subgroups – Test-OD. The same subgroup disparity patterns observed in rMSE persist in MAE, confirming the robustness of the findings across metrics. ModelAllT1DT2DFemaleMale18-3940-6465+ Statistical ARIMA23.5126.2520.7624.2722.7523.0024.9822.58 Naive24.8527.8021.8925.7523.9524.5926.3023.69 AutoETS25.1628.0922.2426.3623.9725.2926.4823.75 ML LightGBM22.1024.4219.7822.9521.2422.0023.2621.06 Linear22.3724.4720.2623.2721.4622.6023.2921.23 XGBoost23.7226.1921.2524.6022.8422.8525.3722.99 CatBoost24.3226.9521.6825.3323.3024.5725.5022.89 Neural NS-Trans.21.9224.1319.7122.6821.1621.5723.2820.94 Transformer22.0124.2719.7522.7221.3021.7323.5220.80 TimeXer22.2324.2820.1923.0121.4621.6423.5421.55 PatchTST22.2524.4720.0222.9121.5821.8023.5821.39 Crossformer22.2624.4720.0523.0321.4921.7523.6121.45 Informer22.2724.3820.1522.9921.5521.5123.6621.67 TFT22.3624.4220.3023.1721.5621.9323.7721.42 MICN22.4024.5620.2523.1721.6421.8623.7721.61 iTransformer22.8125.0620.5623.5022.1222.3723.8922.20 TimesNet22.8525.1120.5823.5022.1922.1024.0122.45 DLinear22.7424.9220.5623.3822.1022.0624.0222.16 SCINet22.7724.8820.6723.5322.0222.1724.0222.15 N-HiTS23.2225.8720.5723.9122.5422.9124.0822.70 TiDE23.2125.4320.9923.9222.5022.3824.5022.77 TSMixer23.3225.7520.9023.8722.7722.6424.5822.77 TimeMixer23.6226.0721.1824.3322.9123.0825.0022.82 N-BEATS24.0426.5321.5624.6723.4223.5325.2823.35 FEDformer28.4131.1125.7129.3327.5027.1429.6428.50 Autoformer32.8235.9029.7434.2831.3530.6134.4633.45 Chronos237.3440.9633.7238.5736.1034.4738.5539.06 API TimeGPT22.7225.2020.2423.8421.6022.4024.2221.56 GPT-5.123.5826.0721.0824.3822.7722.8325.6722.27 GPT-5 Mini–22.4223.01–21.94 Claude 4.524.0726.2921.8524.9723.1724.1025.9822.16 Claude 4.5H24.5227.0821.9625.4323.6124.1726.3023.13 Gemini 327.0129.5224.5028.1025.9125.7129.6325.75 29 AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 10 15 20 25 30 35 30m_rMSE (Mean ± 95% CI) StatsMLNeuralAPI 30m_rMSE by Model (test-id & test-od) test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 0.95 1.00 1.05 OD / ID Ratio OD/ID Ratio by Model (a) rMSE (mg/dL) AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 10 15 20 25 30 35 30m_MAE (Mean ± 95% CI) StatsMLNeuralAPI 30m_MAE by Model (test-id & test-od) test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 0.95 1.00 1.05 OD / ID Ratio OD/ID Ratio by Model (b) MAE (mg/dL) AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 5 10 15 20 25 30m_TIR_error (Mean ± 95% CI) StatsMLNeuralAPI 30m_TIR_error by Model (test-id & test-od) test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 0.95 1.00 1.05 OD / ID Ratio OD/ID Ratio by Model (c) TIR Error (%) Figure 12: 30-minute prediction accuracy: Model comparison between Test-ID and Test-OD splits. Top panel in each subfigure shows metric values with 95% CI; bottom panel shows OD/ID general- ization ratio. 30 AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 15 20 25 30 35 40 1h_rMSE (Mean ± 95% CI) StatsMLNeuralAPI 1h_rMSE by Model (test-id & test-od) test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 0.95 1.00 1.05 OD / ID Ratio OD/ID Ratio by Model (a) rMSE (mg/dL) AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 15 20 25 30 35 1h_MAE (Mean ± 95% CI) StatsMLNeuralAPI 1h_MAE by Model (test-id & test-od) test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 0.95 1.00 1.05 OD / ID Ratio OD/ID Ratio by Model (b) MAE (mg/dL) AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 10 15 20 25 1h_TIR_error (Mean ± 95% CI) StatsMLNeuralAPI 1h_TIR_error by Model (test-id & test-od) test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 0.95 1.00 1.05 OD / ID Ratio OD/ID Ratio by Model (c) TIR Error (%) Figure 13: 1-hour prediction accuracy: Model comparison between Test-ID and Test-OD splits. 31 AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 25.0 27.5 30.0 32.5 35.0 37.5 40.0 2h_rMSE (Mean ± 95% CI) StatsMLNeuralAPI 2h_rMSE by Model (test-id & test-od) test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 0.95 1.00 1.05 OD / ID Ratio OD/ID Ratio by Model (a) rMSE (mg/dL) AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 22.5 25.0 27.5 30.0 32.5 35.0 37.5 2h_MAE (Mean ± 95% CI) StatsMLNeuralAPI 2h_MAE by Model (test-id & test-od) test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 0.95 1.00 1.05 OD / ID Ratio OD/ID Ratio by Model (b) MAE (mg/dL) AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 14 16 18 20 22 24 26 28 30 2h_TIR_error (Mean ± 95% CI) StatsMLNeuralAPI 2h_TIR_error by Model (test-id & test-od) test-id test-od AutoETS Naive ARIMA CatBoost XGBoost Linear LightGBM Chronos2 Autoformer FEDformer N-BEATS TimeMixer TSMixer TiDE N-HiTS SCINet iTransformer DLinear TimesNet MICN TFT Informer Crossformer PatchTST TimeXer Transformer NS-Trans. Gemini 3 Claude 4.5H Claude 4.5 GPT-5 Mini GPT-5.1 TimeGPT 0.95 1.00 1.05 OD / ID Ratio OD/ID Ratio by Model (c) TIR Error (%) Figure 14: 2-hour prediction accuracy: Model comparison between Test-ID and Test-OD splits. 32 18-3940-6465+MFT1DT2D 0 2 4 6 8 10 12 30m rMSE (mg/dL) Fairness: 30m rMSE by Subgroup ARIMA LightGBM PatchTST TFT GPT-5.1 TimeGPT 18-3940-6465+MFT1DT2D 0.8 0.9 1.0 1.1 1.2 OD/ID Ratio (a) 30-minute rMSE (mg/dL) 18-3940-6465+MFT1DT2D 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 1h rMSE (mg/dL) Fairness: 1h rMSE by Subgroup ARIMA LightGBM PatchTST TFT GPT-5.1 TimeGPT 18-3940-6465+MFT1DT2D 0.8 0.9 1.0 1.1 1.2 OD/ID Ratio (b) 1-hour rMSE (mg/dL) 18-3940-6465+MFT1DT2D 0 5 10 15 20 25 30 2h rMSE (mg/dL) Fairness: 2h rMSE by Subgroup ARIMA LightGBM PatchTST TFT GPT-5.1 TimeGPT 18-3940-6465+MFT1DT2D 0.8 0.9 1.0 1.1 1.2 OD/ID Ratio (c) 2-hour rMSE (mg/dL) Figure 15: Fairness comparison across demographic subgroups (age, gender, diabetes type) for all prediction horizons. Top panel in each subfigure shows Test-ID performance; bottom panel shows OD/ID generalization ratio. 33 101520253035 Mean rMSE (mg/dL) 2 3 4 5 6 7 8 9 10 Spread of rMSE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 30m rMSE: Mean vs Spread Stats ML Neural API (a) 30m rMSE 101520253035 Mean MAE (mg/dL) 2 4 6 8 10 Spread of MAE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 30m MAE: Mean vs Spread Stats ML Neural API (b) 30m MAE 20253035 Mean rMSE (mg/dL) 4 5 6 7 8 9 10 Spread of rMSE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 1h rMSE: Mean vs Spread Stats ML Neural API (c) 1h rMSE 1520253035 Mean MAE (mg/dL) 3 4 5 6 7 8 9 Spread of MAE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 1h MAE: Mean vs Spread Stats ML Neural API (d) 1h MAE 2628303234363840 Mean rMSE (mg/dL) 6 7 8 9 10 Spread of rMSE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 2h rMSE: Mean vs Spread Stats ML Neural API (e) 2h rMSE 2224262830323436 Mean MAE (mg/dL) 5 6 7 8 9 Spread of MAE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 2h MAE: Mean vs Spread Stats ML Neural API (f) 2h MAE Figure 16: Accuracy vs. fairness tradeoff across prediction horizons (Spread measure). Columns show rMSE and MAE metrics; rows show 30-minute, 1-hour, and 2-hour predictions. Each point represents a model, colored by family. Lower values on both axes indicate better performance. 34 101520253035 Mean rMSE (mg/dL) 0.030 0.032 0.034 0.036 0.038 0.040 0.042 0.044 0.046 Gini of rMSE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 30m rMSE: Mean vs Gini Stats ML Neural API (a) 30m rMSE 101520253035 Mean MAE (mg/dL) 0.0300 0.0325 0.0350 0.0375 0.0400 0.0425 0.0450 Gini of MAE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 30m MAE: Mean vs Gini Stats ML Neural API (b) 30m MAE 20253035 Mean rMSE (mg/dL) 0.034 0.035 0.036 0.037 0.038 0.039 0.040 0.041 Gini of rMSE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 1h rMSE: Mean vs Gini Stats ML Neural API (c) 1h rMSE 1520253035 Mean MAE (mg/dL) 0.034 0.036 0.038 0.040 0.042 Gini of MAE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 1h MAE: Mean vs Gini Stats ML Neural API (d) 1h MAE 2628303234363840 Mean rMSE (mg/dL) 0.036 0.037 0.038 0.039 0.040 Gini of rMSE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 2h rMSE: Mean vs Gini Stats ML Neural API (e) 2h rMSE 2224262830323436 Mean MAE (mg/dL) 0.035 0.036 0.037 0.038 0.039 0.040 0.041 Gini of MAE Naive ARIMA AutoETS Linear LightGBM CatBoost XGBoost NS-Trans. Transformer TimeXer PatchTST Crossformer Informer TFT MICN TimesNet DLinear iTransformer SCINet N-HiTS TiDE TSMixer TimeMixer N-BEATS FEDformer Autoformer Chronos2 TimeGPT GPT-5.1 Claude 4.5 Claude 4.5H Gemini 3 2h MAE: Mean vs Gini Stats ML Neural API (f) 2h MAE Figure 17: Accuracy vs. fairness tradeoff across prediction horizons (Gini coefficient measure). Columns show rMSE and MAE metrics; rows show 30-minute, 1-hour, and 2-hour predictions. Each point represents a model, colored by family. Lower values on both axes indicate better performance. 35 Experiment Implementation Details Compute Resources All experiments were conducted on a single compute node equipped with 2 NVIDIA RTX 4090 (24GB) GPUs and 128GB RAM. We usedOptunafor hyperparameter tuning and retained the best- performing configurations based on validation performance. Model training was implemented using theNixtlaforecasting library (StatsForecast and MLForecast) for statistical and machine learning models, and the Time-Series-Library (TSLib) framework for 20 neural time-series architectures with fixed hyperparameters and 20-epoch training with early stopping (patience 5). For language model-based experiments, we used API-based inference with deterministic settings. Training time varied depending on model complexity and dataset size. The shallow models, including ARIMA and linear regression, completed training within 1 hour on the FairGlucose splits. Among the deep learning models, N-HiTS was among the most efficient, requiring less than 2 hours in the largest FairGlucose training runs. Transformer-based models such as Transformer and Autoformer were the most computationally intensive, with training times reaching up to 24 hours. All models were trained with early stopping and batch-level validation to ensure computational efficiency without sacrificing performance. Hyperparameters Statistical Models We implemented statistical time series forecasting models using theStatsForecastlibrary [44] and optimized their hyperparameters withOptuna[40]. The optimization aimed to minimize the root mean square error (RMSE) on a validation set by performing a direct hyperparameter search rather than nested auto-model optimization, reducing computational cost substantially. For example, the direct ARIMA optimization takes approximately 30 seconds per trial, compared to 2–5 minutes for ARIMA. The following models were considered in our statistical forecasting framework. ARIMA [30] involved direct optimization of the(p,d,q)and(P,D,Q)parameters. Naive served as a baseline model without tunable hyperparameters. The corresponding hyperparameter search spaces and optimal parameters were defined as follows: • ARIMA: We searched overp ∈ [0, 5],d ∈ [0, 2],q ∈ [0, 5]. Seasonal parameters were disabled (P = D = Q = 0) to focus on non-seasonal patterns. The optimal configuration was: p = 2, d = 1, q = 1. • Naive: This model had no tunable hyperparameters. Machine Learning Forecasting This study uses the Nixtla MLForecast framework [39] to implement machine-learning models for time-series forecasting. MLForecast is a scalable, production-grade library that automatically converts time-series data into supervised learning features (lags, rolling statistics, date features), then fits any estimator following thescikit-learnAPI (havingfit/predictmethods). The library handles recursive forecasting and feature updates over the prediction horizon, and supports horizontal scaling viapandas,polars,Dask,Spark, orRayexecution [39].MLForecastautomates much of the feature engineering process: it generates lag features, applies lag-based transformations, and extracts relevant date-based features. The user only needs to specify the desired regressors and configuration options, while the library manages updates and iterations over time steps internally. The models evaluated in this work were selected from standard regression algorithms compatible with MLForecast:LinearRegression(ordinary least squares) [43],lightgbm.LGBMRegressor[18], xgboost.XGBRegressor[32], andcatboost.CatBoostRegressor[31]. These are instantiated via MLForecast by passing a list of regressor instances. MLForecast then wraps them into a forecasting pipeline with lag-feature and date-feature generation. For each model, we defined a model-specific Optuna hyperparameter search space (see the code listing), including, e.g.,n_estimators,max_depth,learning_ratefor boosting;l2_leaf_reg, 36 border_countfor CatBoost; andfit_interceptfor LinearRegression. The objective function trains MLForecast with the current trial’s parameters on the training split (using the full set of lag features up toINPUT_LENGTH) and evaluates on the validation split using recursive multi-step forecasts and average root-mean-squared error (rMSE). A fixed lag vector oflags = [1, 2, ..., INPUT_LENGTH] is used without additional lag transforms. Hyperparameter search spaces and optimal parameters used in this work: • LinearRegression:fit_intercept ∈ True, Falseandn_jobs=−1. The optimal configuration was: fit_intercept = True, n_jobs =−1. •LightGBM:n_estimators ∈[50, 500](step 50);max_depth ∈[3, 15]; learning_rate ∈[0.01, 0.3];num_leaves ∈[10, 300](step10); min_child_samples∈ [5, 100](step 5);subsample∈ [0.5, 1.0];colsample_bytree∈ [0.5, 1.0]; reg_alpha∈ [0.0, 10.0]; reg_lambda∈ [0.0, 10.0]. The optimal configuration was:n_estimators= 300,max_depth= 10,learning_rate = 0.04,num_leaves= 180,min_child_samples= 75,subsample= 0.7, colsample_bytree = 0.5, reg_alpha = 6.5, and reg_lambda = 4.5. • XGBoost: same as LightGBM plusgamma∈ [0.0, 5.0]andmin_child_weight∈ [1, 10]. The optimal configuration was:n_estimators= 50,max_depth= 5,learning_rate= 0.2,subsample= 0.9,colsample_bytree= 0.7,reg_alpha= 4.25,reg_lambda= 3.5, gamma = 0.6, and min_child_weight = 7. •CatBoost:iterations∈ [50, 500](step 50);depth∈ [3, 10](step 1);learning_rate∈ [0.01, 0.3] (step 0.01);l2_leaf_reg∈ [0.1, 10.0](step 0.5);border_count∈ [32, 255] (step 32). The optimal configuration was:iterations= 350,depth= 3,learning_rate = 0.28, l2_leaf_reg = 8.1, and border_count = 32. All models were trained using recursive forecasts over the horizon (OUTPUT_LENGTH), matching the temporal frequency (freq=5min). Optuna optimization was run with a fixed number of 50 trials, using the validation rMSE as the objective. The final model for each best trial was refit using MLForecast with the optimal parameter set, yielding a MLForecast model object ready for out-of-sample forecasting and evaluation. Neural Time-Series Models (TSLib) We trained 20 neural time-series architectures using the Time-Series-Library (TSLib) framework, which provides a unified training pipeline across diverse architectures. All models use fixed hyperpa- rameters with 20-epoch training and early stopping (patience 5), input length 288 (24 hours), and prediction length 24 for the main 2-hour evaluation. The architectures include attention-based models (PatchTST, iTransformer, TimeXer, Crossformer, NS-Transformer, Transformer, Informer, Auto- former, FEDformer, TFT), MLP/linear models (DLinear, TiDE, TSMixer, TimeMixer, N-BEATS, N-HiTS, MICN), convolutional models (TimesNet, SCINet), and the zero-shot foundation model Chronos2. Below we describe the hyperparameter configurations for representative architectures. PatchTST (PatchTST) [17]:An efficient channel-independent transformer that uses se- quence patching for scalable long-horizon forecasts [17]. The Optuna search space included: patch_len ∈ [8, 32](step 4);stride ∈ [4, 16](step 4);hidden_size ∈ [64, 512](step 32);n_heads ∈ 2, 4, 8, 16, 32;dropout ∈ [0.00, 0.30](step 0.05);learning_rate ∈ 1e −6 , 5e −6 , 1e −5 ,..., 1e −2 ;batch_size ∈ 32, 64, 128, 256; andmax_stepsderived from 10–25 epochs. The optimal configuration was:patch_len= 8,stride= 8,hidden_size= 320, n_heads = 8, dropout = 0.1, learning_rate = 2e −4 , batch_size = 256. DLinear (DLinear) [37]: A linear decomposition model with separate trend and seasonal- ity paths, enabling quick convergence for long sequences [37]. The Optuna search space in- cluded:learning_rate∈1e −6 , 5e −6 , 1e −5 ,..., 1e −2 ;batch_size∈32, 64, 128, 256; and max_stepsderived from 3–10 epochs. The optimal configuration was:learning_rate=1e −3 , batch_size = 64. 37 Transformer (VanillaTransformer) [33]: Standard encoder–decoder transformer with multi- head attention and an MLP decoder, serving as a baseline for modeling long-range dependencies [33]. The Optuna search space included:hidden_size∈ [32, 128]with a step of32;n_head∈4, 6; encoder_layers ∈ 2, 3, 4;decoder_layers ∈ 2, 3, 4;dropout ∈ [0.00, 0.30]with a step of0.05;learning_rate ∈ 1e −6 ,..., 1e −2 ;batch_size ∈ 32, 64; andmax_steps derived from 5–10 epochs. The optimal configuration was:hidden_size= 128,n_head= 4,encoder_layers= 4,decoder_layers= 2,dropout= 0.15,learning_rate=5e −4 , batch_size = 64, activation = gelu. NHITS (NHITS) [38]:A hierarchical interpolation model that divides forecasting into frequency-specialized submodels to improve accuracy and efficiency [38].Optuna search space:n_pool_kernel_size ∈[2,2,2],[4,4,4];n_freq_downsample ∈[8,4,1], [16,8,1];interpolation_mode ∈linear,nearest,cubic;n_blocks ∈[1,1,1], [2,2,2];mlp_units ∈[[512,512],...],[[256,256],...];activation ∈ReLU, Softplus,Tanh; along withlearning_rate,batch_size, andmax_stepsderived from 8–20 epochs.The optimal configuration was:n_pool_kernel_size=[2,2,2], n_freq_downsample=[16,8,1],interpolation_mode=nearest,n_blocks=[2,2,2], mlp_units=[[512,512],[512,512],[512,512]],activation=ReLU,learning_rate= 1e −5 , batch_size = 256. TFT (TFT) [36]: Temporal Fusion Transformer, which incorporates gating mechanisms, recur- rent encoders, and a multi-head attention decoder for interpretable forecasting and variable se- lection. The Optuna search space included:hidden_size ∈ 32, 64, 128, 256;n_rnn_layers ∈ 1, 2, 3;dropout ∈ [0.00, 0.30]with a step of0.05; along withlearning_rate,batch_size, andmax_stepscorresponding to 5–15 epochs. The optimal configuration was:hidden_size= 64, n_rnn_layers = 3, dropout = 0.1, learning_rate = 1e −4 , and batch_size = 64. Autoformer (Autoformer) [35]: A decomposition transformer that integrates progressive trend/seasonal splitting and auto-correlation mechanisms for efficient periodic pattern modeling. Op- tuna search space:encoder_layers∈1, 2, 3, 4, 5, 6;decoder_layers∈1, 2;hidden_size ∈32, 64, 128, 256;n_head ∈2, 4, 8;conv_hidden_size ∈16, 32, 64;activation ∈ gelu,relu;dropout∈[0.10, 0.40] step 0.05; pluslearning_rate,batch_size,max_steps derived from 8–20 epochs. The optimal configuration was:hidden_size= 32,encoder_layers= 3,decoder_layers= 2,n_head= 2,conv_hidden_size= 32,activation=relu,dropout= 0.4, learning_rate = 5e −3 , batch_size = 256. N-BEATS (NBEATS) [45]: A neural basis expansion model with interpretable forecast and back- cast stacks for univariate time-series forecasting. The architecture uses doubly residual stacking of blocks to model trend and seasonality components separately. Optuna search space:stack_types ∈identity,trend,seasonality;n_blocks ∈1, 2, 3, 4, 5;mlp_unitsconstructed from two sampled integers∈[64, 512] with step 64; along withlearning_rate,batch_size, andmax_stepsderived from 8–20 epochs. The optimal configuration was:stack_types =[seasonality],n_blocks=[5],mlp_units=[[448, 512]],learning_rate=1e −4 , batch_size = 32. LLM for Glucose Forecasting We implement a prompt-based glucose prediction agent using large language models (LLMs) to generate short-term glucose forecasts from historical CGM sequences. Given a sequence of CGM valuesx 1 ,x 2 ,...,x T sampled every 5 minutes and a desired prediction horizonH(defined as the number of future 5-minute intervals), the agent formats the input into a natural language prompt that encodes domain knowledge and physiological constraints. This prompt is then passed to a specified LLM backend using LangChain’s unified interface. The model’s textual output, expected to be in a strict JSON format, is parsed to extract the predicted future glucose valuesˆx T+1 ,..., ˆx T+H . The prompt guides the model to return predictions that are physiologically plausible, including glucose values between 30–400 mg/dL, with typical rates of change bounded between−5and+5 mg/dL per minute. It also encourages the model to account for known glucose dynamics such as 38 glucose inertia and time-of-day effects (e.g., dawn phenomenon and meal-related spikes). Predictions that do not conform to these structural or physiological constraints are flagged for review or discarded. The agent also includes a fallback interface to the Nixtla time-series API, which provides non- LLM-based forecasting using classical methods. All models follow the same prompt and output specification to ensure comparability across different model families and implementations. Prompt Template The prompt issued to the LLMs is shown below: Implementation Details We implement an end-to-end evaluation pipeline to benchmark LLM-based models for glucose forecasting using the FairGlucose dataset. The system is modular, supports multiple backends, and operates on a standardized interface for CGM input handling, model inference, and evaluation. The pipeline is orchestrated through a Python script that accepts various command-line arguments, including the model type, model name, task index, number of samples, input length, and output horizon. These parameters allow for parallel and distributed execution over patient records using task batching. Model access credentials (e.g., OpenAI, Anthropic, Google Vertex AI) are dynamically configured via environment variables, allowing authenticated API access at runtime. Each selected CGM record is sliced into an input window of 24 hours (288 values sampled every 5 minutes) and a prediction horizon of up to 8 hours (3–96 steps), depending on the experimental setup. The main LLM comparison uses a 24-step, 2-hour prediction horizon. Glucose Prediction Prompt System Prompt: You are a precise time-series forecasting assistant for glucose prediction. You must return ONLY a JSON array of integers (predicted glucose values in mg/dL). Example format: [120, 125, 130, 128, 126, 124, 122, 120, ...] Do not include any explanatory text, markdown formatting, or additional fields. Return ONLY the JSON array of integers. User Prompt: You are a glucose prediction model for diabetes patients. Given historical CGM (Continuous Glucose Monitor) readings at 5-minute intervals, predict the next forecast_horizon glucose values. Historical CGM values (most recent input_size readings, in mg/dL): cgm_history Instructions: 1. Analyze the glucose trends, patterns, and rate of change 2. Consider typical post-meal curves and insulin effects 3. Predict the next forecast_horizon glucose values (5-minute intervals) 4.Return ONLY a JSON array of forecast_horizon numbers (predicted glucose values in mg/dL) Example output format: [120, 125, 130, 128, 126, 124, 122, 120, 118, 116, 115, 114, ...] Your prediction (JSON array only): For each record, the system constructs a structured prompt and passes it to the selected LLM using theGlucosePredictoragent. The agent handles the model-specific API interaction, parsing, and prediction recovery. If the selected model is a time-series backend (e.g., Nixtla’s TimeGPT [19]), the system constructs a timestamped time series and passes it directly to the statistical model API. Parallel processing using a thread pool is used to accelerate batch inference. Predictions are stored in a structured format that includes the patient ID, prediction time, input sequence, ground truth, and model prediction. After inference, the results are saved in Parquet format and evaluated using a custom sequence evaluation toolkit. This toolkit calculates root mean 39 square error (rMSE) over multiple horizons and supports subgroup-level evaluation (e.g., by diabetes type, age interval, or prediction hour). The evaluation framework ensures consistency across all experiments and provides both neat and full-format reports for downstream analysis. The key parameters for LLM-based glucose forecasting are: input context window of 288 time steps (24 hours at 5-minute intervals), forecast horizon of 24 time steps (2 hours ahead), and temperature set to 0.0 for deterministic generation. All LLM models (GPT-5.1, GPT-5 Mini, Claude 4.5 Sonnet, Claude 4.5 Haiku, Gemini 3 Flash) use the same prompt template and configuration to ensure fair comparison. The system supports a wide range of model backends, including GPT-5.1, GPT-5 Mini, Claude 4.5 Sonnet, Claude 4.5 Haiku, Gemini 3 Flash, and TimeGPT [19]. Each model is evaluated under the same input-output specification and metrics to ensure fair comparison. The resulting predictions and performance metrics are used to assess fairness, subgroup robustness, and generalization across temporal and demographic dimensions. 40