Paper deep dive
Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN
Eliuvish Han Cui
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Anti-amyloid therapies and blood-based biomarkers are changing Alzheimer disease workups into a two-stage measurement workflow: screen broadly with cheaper information, then spend scarce confirmatory amyloid measurements where they support the decision that will be reported. Amyloid positron-emission tomography (PET) remains one such protocol measurement for amyloid burden, but PET slots, trial budgets, and payer-facing evidence packages are finite. This paper asks a deliberately operational question: when is simple transparent PET validation enough, and when is a fitted residual-uncertainty score worth the added complexity? For a weighted protocol target, the first-order value of validating subject i is the product of target influence and residual protocol uncertainty. Generic uncertainty sampling uses only the second factor and can spend PET measurements on subjects that are hard to predict but weak for the scientific, clinical, or commercial claim. We apply this rule to the A4/LEARN PET archive, treating observed PET as a design laboratory for scarce-confirmation studies. For the primary APOE4 carrier versus non-carrier contrast in Centiloid 24-or-higher PET positivity, simple APOE4-balanced validation recovers nearly all of the target-specific gain: at PET budget 200, the confidence-interval width ratio relative to random validation is 0.923 for APOE4 balancing and 0.914 for target-specific scoring, while generic uncertainty sampling is 0.980. Other targets behave differently: target-specific scoring gives larger gains for an age-slope analysis and for cutoff-indexed PET positivity. The practical message is simple: spend scarce protocol measurements according to the claim being validated, not only according to prediction uncertainty.
Tags
Links
- Source: https://arxiv.org/abs/2608.22223v1
- Canonical: https://arxiv.org/abs/2608.22223v1
Trouble viewing inline? Open PDF directly →
Full Text
28,622 characters extracted from source content.
Expand or collapse full text
Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN Elvis Han Cui Clinical Functional Service Provider (cFSP) Department Kuntuo, an IQVIA company Department of Biostatistics, University of California, Los Angeles August 23, 2026 Abstract Anti-amyloid therapies and blood-based biomarkers are changing Alzheimer disease workups into a two-stage measurement workflow: screen broadly with cheaper information, then spend scarce confirmatory amyloid measurements where they support the decision that will be re- ported. Amyloid positron-emission tomography (PET) remains one such protocol measurement for amyloid burden, but PET slots, trial budgets, and payer-facing evidence packages are finite. This paper asks a deliberately operational question: when is simple transparent PET validation enough, and when is a fitted residual-uncertainty score worth the added complexity? For a weighted protocol target, the first-order value of validating subject i is the product of target influence and residual protocol uncertainty. Generic uncertainty sampling uses only the second factor and can spend PET measurements on subjects that are hard to predict but weak for the scientific, clinical, or commercial claim. We apply this rule to the A4/LEARN PET archive, treating observed PET as a design laboratory for scarce-confirmation studies. For the primary APOE4 carrier versus non-carrier contrast in Centiloid 24-or-higher PET positivity, simple APOE4-balanced validation recovers nearly all of the target-specific gain: at PET budget 200, the confidence-interval width ratio relative to random validation is 0.923 for APOE4 balancing and 0.914 for target-specific scoring, while generic uncertainty sampling is 0.980. Other targets behave differently: target-specific scoring gives larger gains for an age-slope analysis and for cutoff-indexed PET positivity. The practical message is simple: spend scarce protocol measure- ments according to the claim being validated, not only according to prediction uncertainty. Keywords: amyloid PET; active sampling; design-based inference; two-phase validation; APOE4. 1 Why scarce confirmatory measurement is the problem Validation studies often begin with an awkward asymmetry. A large archive contains variables that are cheap enough to collect broadly, while the measurement that defines the scientific endpoint is expensive, invasive, slow, or restricted by assay capacity. The operational question is not merely 1 arXiv:2608.22223v1 [stat.AP] 23 Aug 2026 whether a surrogate model predicts well. It is which expensive confirmations should be bought, scanned, assayed, or adjudicated when the validation budget is smaller than the screened population. This question has become commercially and clinically sharper in Alzheimer disease. Anti- amyloid therapies require evidence of amyloid pathology before treatment, making biomarker con- firmation part of the treatment workflow rather than only a research endpoint [U.S. Food and Drug Administration, 2026, 2024]. At the same time, blood-based biomarkers and other lower-cost mea- sures are expanding the front end of the diagnostic funnel [Palmqvist et al., 2024, 2025]. Expanded coverage and broader screening do not eliminate the PET bottleneck; they make the allocation problem more visible [Centers for Medicare & Medicaid Services, 2023]. A diagnostic company, trial sponsor, imaging network, or payer-facing evidence team may care about an APOE4 contrast, an age trend, a cutoff choice, or a treatment-eligibility rule. Those are different targets, and they need not be served by the same validation subset. Amyloid PET in A4/LEARN makes the distinction concrete. The Anti-Amyloid Treatment in Asymptomatic Alzheimer’s Disease study and the companion LEARN study contain genetic, de- mographic, cognitive, plasma, and magnetic-resonance imaging information, together with amyloid PET on a large subset of participants [Sperling et al., 2020, 2023]. PET is costly but scientifically central: it is the protocol measurement for amyloid burden, summarized here as Centiloid 24-or- higher positivity [Clark et al., 2012, Klunk et al., 2015, Navitsky et al., 2018, Jack et al., 2024]. In a future study, investigators might screen broadly with cheaper information and confirm PET on only a subset. Which subset should be validated? The answer depends on the question the PET measurements are meant to answer. The tempting response is to build the best possible PET-risk model and validate subjects whose outcomes are most uncertain. That response is natural if the business objective is to improve the predictor. It is not always natural if the objective is to support a prespecified claim, contrast, or regulatory-style evidence package. This paper focuses on the APOE epsilon-4 (APOE4) carrier versus non-carrier contrast in PET positivity. APOE4 is measured before PET, is biologically established as an Alzheimer disease risk factor, and defines a clinically interpretable target contrast [Farrer et al., 1997, Sperling et al., 2020]. For this target, a transparent APOE4-balanced PET validation design performs almost as well as a fitted target-specific design. That is the main empirical lesson. The secondary lesson is just as important: the same conclusion does not automatically transfer to an age slope, a cutoff-indexed PET endpoint, or a cohort-structured diagnostic contrast. The contribution is therefore not another claim that a surrogate model predicts PET well. The contribution is a decision rule for using surrogate information without letting the surrogate redefine the target. The rule is old in spirit, drawing on two-phase sampling, model-assisted survey estimation, and missing-by-design inference [Neyman, 1934, Cochran, 1977, Särndal et al., 1992, Robins et al., 1994, Breslow et al., 2009, Gilbert et al., 2014]. The reframing for this application is operational: define the claim, identify the protocol measurement that anchors it, record validation probabilities before outcome use, and spend scarce measurements where they reduce uncertainty for that claim. The paper can be read in four steps: Section 2 defines the A4/LEARN scarce- 2 PET decision, Section 3 gives the recorded-probability rule, Section 4 shows when simple APOE4 balancing is enough, and Sections 5–6 show when target-specific scoring or pilot learning becomes more useful. 2 The A4/LEARN validation decision The analysis uses the July 2024 A4/LEARN controlled-access release. The processed first-phase archive contains 6,945 rows; 4,492 have nonmissing Centiloid PET; and 4,460 have both PET and APOE4 status. Individual-level data are not redistributed. The complete PET archive is used retrospectively as a design laboratory: hide PET outcomes, spend an emulated PET validation budget according to candidate rules, and evaluate finite-archive PET inference for prespecified targets. The primary endpoint is Y i = 1C i ≥ 24, where C i is Centiloid PET. CL24 is used as an operational florbetapir/Centiloid protocol threshold, not as a universal disease boundary. The primary target is the finite-archive APOE4 carrier versus non-carrier difference in CL24 positivity over the observed design-emulation archive; population language is shorthand for this empirical contrast, not an external-validity claim. The surrogate score uses first-phase demographic, APOE4, cognitive, plasma, and MRI informa- tion where supported. Scores are built and evaluated under cross-fitting when PET outcomes are used to assess prediction, so an individual’s own PET value is not fed back into its predicted risk. Cross-fitting, however, is not the same thing as prospective availability. Demographics, APOE4, and design variables have broad first-phase support; cognitive summaries are available for 1,708 PET+APOE participants with at least one listed variable and 1,167 with all listed variables; plasma summaries are available for 1,926 and 1,349 under the same convention; pTau217 has 1,439 and 1,000; and MRI summaries have 1,792 complete cases. The main claims therefore concern validation design under documented support, not clinical deployment of a particular PET-risk model. The design comparison includes four practically distinct choices: Random validationbaseline design-based sampling; Generic uncertaintyvalidate subjects with high PET prediction uncertainty; APOE4-balanced validation spend equal expected PET counts in carrier strata; Target-specific validationallocate by target influence times residual uncertainty. All ratios below are relative to random validation at the same PET budget. Budget 200 is a readable mid-range scarce-PET scenario in a broader 50/100/200/400/800 budget curve, not a uniquely optimized choice. Oracle lower-bound rows are used only as infeasible diagnostics and are not proposed as deployable designs. The intervals are finite-archive design intervals for validation randomness; they do not add superpopulation sampling variation, surrogate-model selection uncer- tainty, future feature-availability uncertainty, or biological measurement-error components beyond 3 the chosen PET protocol endpoint. A Recorded-probability validation design first-phase information W i surrogate score i residual uncertainty i target weight |a i | target value |a i (t)| i (t) finite-archive design interval PET residual correction record i before PET 100200400800 Total PET budget 0.04 0.06 0.08 0.10 0.12 RMSE B. Estimation error 100200400800 Total PET budget 0.05 0.10 0.15 0.20 Mean 95% CI half-width C. Interval width 100200400800 Total PET budget 0.88 0.90 0.92 0.94 0.96 0.98 Empirical coverage D. Coverage Random Uncertainty Pilot target-specific Pilot target + random Oracle lower bound Figure 1: Recorded-probability validation workflow. First-phase information supports prediction and residual-uncertainty scoring; validation probabilities are recorded before PET outcomes enter the final correction; inference reports PET-defined design intervals. 3 Target-aware rule and recorded estimator Let Y i (t) be the protocol measurement for subject i, indexed by threshold or target label t, and let θ(t) = X i a i (t)Y i (t) be the finite-archive target. The weight a i (t) encodes what the study is trying to estimate. In an APOE4 contrast, it separates carriers from non-carriers; in an age-slope analysis, it reflects age leverage; in a cutoff curve, it changes with the PET threshold. Let μ i (t) be a first-phase prediction of the PET-defined quantity and let σ i (t) describe residual protocol uncertainty after the first-phase information is used. A validation design records a positive probability π i before PET outcomes enter the final correction. The augmented estimator is b θ(t) = X i a i (t)μ i (t) + X i R i π i a i (t)Y i (t)− μ i (t), 4 where R i indicates validation. Principle 1 (Target-aligned validation value). For a scalar target, the first-order value of validating subject i is proportional to |a i (t)|σ i (t). For an indexed target, replace this by a prespecified summary such as S i = ( X t ω t a i (t) 2 σ i (t) 2 ) 1/2 . This is the design message. Generic uncertainty sampling uses σ i (t). Target-aligned valida- tion uses |a i (t)|σ i (t). The difference matters whenever prediction uncertainty and target influence point to different subjects. Conditional on the archive, the first-phase scores, and the recorded probabilities, E R " X i R i π i − 1 a i (t)Y i (t)− μ i (t) # = 0. Thus the surrogate model can improve precision, but correctness of the surrogate model is not what centers the estimator. Centering comes from the recorded validation probabilities and positivity. In a real deployment, S i must be computed from first-phase, historical, cross-fitted, or pilot information available before the final PET validation decision. Full-PET residuals are used here only to evaluate candidate designs in the retrospective laboratory. 4 Primary APOE4 result: simple balancing is nearly enough For the APOE4 contrast, the PET protocol contrast is 0.334 and the surrogate-only contrast is about 0.355. The cross-fitted surrogate AUC is about 0.779, but the AUC is not the design conclusion. The design conclusion is that uncertainty sampling adds little for this estimand, while APOE4 balancing recovers most of the target-specific gain. At PET budget 200, the CI-width ratio is 0.980 for generic uncertainty sampling, 0.923 for APOE4-balanced validation, and 0.914 for target-specific scoring. In words, the fitted target-specific design is best among these rules, but the simple balanced design is almost indistinguishable for practical planning. Balanced validation is easier to explain, audit, and implement than a residual- uncertainty score, so it should be the serious default comparator for this target. 5 APOE4 carrier contrast Age slopeCutoff-indexed APOE4 0.80 0.85 0.90 0.95 1.00 1.05 CI half-width / random validation 0.980 1.000 0.905 0.923 1.001 0.914 0.806 0.859 Validation-design efficiency at PET budget 200 Generic uncertaintySimple balanced comparatorTarget-specific Figure 2: Validation-design efficiency at PET budget 200. The APOE4 result is the primary planning result: simple APOE4 balancing nearly matches target-specific scoring. The age and cutoff rows show the scope condition: target-specific residual-uncertainty scoring is more useful when influence is heterogeneous or the target is indexed. This result should not be oversold. It does not say that fitted scores are unnecessary in every PET validation study. It says that, for this two-group target, the most important influence structure is already visible before modeling: APOE4 status defines the contrast. Once the design guaran- tees adequate PET measurement in both strata, the remaining target-specific residual-uncertainty refinement is modest. A useful stress check removes APOE4 from the PET-risk surrogate while keeping APOE4 as the prespecified target-defining variable. That stricter score has AUC 0.724 and compresses the APOE4 surrogate-only contrast to 0.069 versus the PET contrast 0.334, giving bias -0.265. Target-specific validation still improves precision relative to random validation at budget 200, with CI-width ratio 0.952 and RMSE ratio 0.961; generic uncertainty sampling has CI-width ratio 1.004. This separates APOE4’s role as an estimand-defining stratum from its optional role inside a prediction model. Table 1: Design ratios at PET budget 200 for selected A4/LEARN targets. Ratios are relative to random validation; values below 1 indicate narrower intervals. One-dimensional balanced rows are checks, not a universal dominance ranking. TargetGeneric uncertainty Balanced comparator Target-specific APOE4 carrier contrast0.980–0.9880.9230.914–0.915 Age slope1.0001.0010.806 CL24 within cutoff-indexed APOE4 curve0.905–0.859 6 5 Scope checks: when targeting matters The APOE4 result is deliberately paired with scope checks. For the age slope in PET positivity, the surrogate-only point estimate is close to the PET protocol value, but target-specific validation still matters for precision: the CI-width ratio at budget 200 is 0.806, while generic uncertainty sampling is 1.000 and age-quartile balancing is 1.001. Coarse balancing is not the same as allocating by leverage and residual protocol uncertainty. For cutoff-indexed APOE4 positivity over Centiloid thresholds 20 to 30, the target changes with the threshold. At the primary threshold c = 24, the PET contrast is 0.334 and the threshold-specific surrogate contrast is 0.347. A cutoff-range design gives a CI-width ratio of about 0.859, compared with about 0.905 for generic uncertainty validation. The endpoint is not a single-cutoff artifact: at Centiloid thresholds 20, 24, and 30, the PET APOE4 contrasts are 0.326, 0.334, and 0.318. 202224262830 Centiloid threshold 0.320 0.325 0.330 0.335 0.340 0.345 0.350 APOE4+ minus APOE4- CL 24 A APOE4 PET cutoff sensitivity PET protocol Surrogate-only 202224262830 Centiloid threshold 0.000 0.005 0.010 0.015 0.020 0.025 0.030 Surrogate-only minus PET B Surrogate bias across cutoffs 202224262830 Centiloid threshold 0.80 0.85 0.90 0.95 1.00 CI half-width / random C Validation efficiency, budget 200 Uncertainty Fixed CL24 Integrated cutoff Operational PET cutoff sensitivity, with process-band details reserved for supplement Figure 3: APOE4 PET cutoff sensitivity across operational Centiloid positivity thresholds. The target-specific design is evaluated over the prespecified threshold range rather than only at CL24. The A4-versus-LEARN comparison is moved to the diagnostic layer. It is not a causal or biological cohort contrast. Its value is to show that surrogate error can align with cohort structure: a cohort-unaware surrogate has AUC about 0.779 but gives a surrogate-only A4-versus-LEARN contrast of 0.380 compared with the PET protocol contrast of 0.873. Adding cohort information raises AUC to about 0.908 and reduces the contrast-aligned bias from -0.493 to -0.039. This stress test supports the target-aware warning but should not occupy the primary biomedical story. 6 Deployable pilot evidence The retrospective archive can use all PET outcomes to evaluate designs, but a real study cannot. The deployable version is a two-wave design: random PET pilot−→ learn residual uncertainty −→ record target-aligned probabilities −→ estimate with the recorded design. 7 In the existing A4/LEARN pilot emulation, a random pilot of 50 followed by a target-specific wave of 150 gives an RMSE ratio of 0.844 and a CI-width ratio of 0.919 relative to pilot-random validation, with coverage 0.954. At total budget 400 with a random pilot of 100, the corresponding ratios are 0.917 and 0.918, with coverage 0.950. These numbers are less dramatic than full-archive oracle diagnostics, but they are more relevant for practice because they do not require unobserved PET outcomes to choose the second wave. The pilot design also clarifies the role of machine learning. A richer learner may reduce prediction error, but it is useful for validation only if it improves residual uncertainty in the influential part of the target and is available before the validation decision. In the model-ladder audit, flexible safe-feature scores support the same qualitative APOE4 conclusion, while the pTau217 sensitivity is kept outside the main design because its archive timing is not uniformly pre-PET. 7 Discussion The paper’s message is intentionally narrower than the earlier technical version. A good surrogate model is useful, but validation should be aligned with the target that PET is meant to estimate. In A4/LEARN, this distinction leads to a practical planning rule. For the primary APOE4 PET con- trast, start with APOE4-balanced validation; it is transparent and nearly matches the target-specific design. For targets with uneven influence, cutoff-indexed structure, or within-stratum residual het- erogeneity, move to target-aligned residual-uncertainty scoring. For prospective deployment, learn residual uncertainty from historical data or a random pilot and record validation probabilities before PET outcome use. Several boundaries remain. The analysis is retrospective: the complete PET archive is used to evaluate scarce-validation designs, not to claim that future PET outcomes would be observed under the same mechanism. PET visual read and CL24 agree in 92.2% of 4,492 paired rows, with Cohen’s kappa 0.821, but visual read is a plausibility audit rather than a replacement endpoint. Plasma, MRI, pTau217, and cognitive features have heterogeneous timing and support, so they should be used for targeting only when available in the intended first-phase cohort. Crossed stratified grids and cutoff-aware balanced validation remain outside the current performance package. Those limitations are also why the design convention matters. When expensive protocol mea- surements define the endpoint, the study should document the target, the first-phase support, the validation probabilities, and the final PET-defined estimator. Prediction can help decide where to measure, but the measurement design should remain answerable to the scientific estimand. The same discipline is useful beyond this PET example, but it should not be forced into the main empirical claim. In biophysical chemistry and stochastic biochemical kinetics, the scientific object is often a kinetic, thermodynamic, or trajectory-level quantity, while cheaper proxy measurements may be available before a high-fidelity assay or imaging protocol is run. Work associated with Qian’s stochastic thermodynamic view of biochemical systems is a useful conceptual reminder: the measured trajectory, state variable, flux, or protocol-defined functional has to be specified before 8 auxiliary measurements are allowed to guide inference [Qian, 2006, 2007]. The present paper does not model biochemical reaction networks; it borrows only the measurement-design lesson that proxy information should guide protocol measurement, not replace the protocol quantity. A separate mathematical extension would replace the binary PET target by a time-indexed or paired-time endpoint. Product-integral and bivariate survival functionals, including Dabrowska- type estimators on the plane, give a natural language for such targets [Dabrowska, 1988, Andersen et al., 1993, Hougaard, 2000]. That extension is compatible with the recorded-probability contract, but it is not needed for the A4/LEARN empirical message. Keeping it as an extension makes the paper easier to read: the current claim is about scarce confirmatory PET measurements for scalar and cutoff-indexed amyloid targets. Data availability Individual A4/LEARN participant-level data are controlled-access study data available to qual- ified researchers through the A4 and LEARN Study Data Portal, A4StudyData.org, subject to registration, approval, and portal data-use requirements. Controlled-access source files are not redistributed. The accompanying reproducibility materials provide manuscript sources, analysis and simulation code, cleared aggregate tables and figures, requirements information, and synthetic simulation checks. Use of AI/NLP tools OpenAI Codex and ChatGPT were used as editorial and programming assistants for drafting, code organization, formatting checks, and submission preparation. The author reviewed and takes responsibility for all manuscript content, code, analyses, and conclusions. References Per Kragh Andersen, Ornulf Borgan, Richard D. Gill, and Niels Keiding. Statistical Models Based on Counting Processes. Springer, New York, 1993. Norman E. Breslow, Thomas Lumley, Christie M. Ballantyne, Lloyd E. Chambless, and Michal Kulich. Improved horvitz-thompson estimation of model parameters from two-phase stratified samples. Statistics in Biosciences, 1(1):32–49, 2009. doi: 10.1007/s12561-009-9001-6. Centers for Medicare & Medicaid Services. Beta amyloid positron emission tomogra- phy in dementia and neurodegenerative disease. National Coverage Analysis decision memo CAG-00431R, 2023. URL https://w.cms.gov/medicare-coverage-database/view/ ncacal-decision-memo.aspx?NCAId=308. Accessed 2026-08-23. 9 Christopher M. Clark, Michael J. Pontecorvo, Thomas G. Beach, Barry J. Bedell, R. Edward Coleman, P. Murali Doraiswamy, Adam S. Fleisher, Eric M. Reiman, Marwan N. Sabbagh, Carl H. Sadowsky, Julie A. Schneider, Mark A. Mintun, Daniel M. Skovronsky, et al. Cerebral PET with florbetapir compared with neuropathology at autopsy for detection of neuritic amyloid- beta plaques: A prospective cohort study. The Lancet Neurology, 11(8):669–678, 2012. doi: 10.1016/S1474-4422(12)70142-4. William G. Cochran. Sampling Techniques. John Wiley & Sons, 3 edition, 1977. Dorota M. Dabrowska. Kaplan-meier estimate on the plane. The Annals of Statistics, 16(4):1475– 1489, 1988. doi: 10.1214/aos/1176351049. Lindsay A. Farrer, L. Adrienne Cupples, Jonathan L. Haines, Bradley Hyman, Walter A. Kukull, Richard Mayeux, Richard H. Myers, Margaret A. Pericak-Vance, Neil Risch, and Cornelia M. van Duijn. Effects of age, sex, and ethnicity on the association between apolipoprotein e genotype and alzheimer disease: A meta-analysis. JAMA, 278(16):1349–1356, 1997. doi: 10.1001/jama. 1997.03550160069041. Peter B. Gilbert, Xiaoying Yu, and Andrea Rotnitzky. Optimal auxiliary-covariate-based two- phase sampling design for semiparametric efficient estimation of a mean or mean difference, with application to clinical trials. Statistics in Medicine, 33(6):901–917, 2014. doi: 10.1002/sim.6006. Philip Hougaard. Analysis of Multivariate Survival Data. Springer, New York, 2000. Clifford R. Jack, J. Scott Andrews, Thomas G. Beach, Teresa Buracchio, Billy Dunn, Matthew Graf, Oskar Hansson, Clifford Ho, William Jagust, Eric McDade, Jose Luis Molinuevo, Gil D. Rabinovici, Christopher C. Rowe, Reisa Sperling, Charlotte Teunissen, Maria C. Carrillo, et al. Revised criteria for diagnosis and staging of alzheimer’s disease: Alzheimer’s association work- group. Alzheimer’s & Dementia, 20:5143–5169, 2024. doi: 10.1002/alz.13859. William E. Klunk, Robert A. Koeppe, Julie C. Price, Tammie L. Benzinger, Michael D. Devous, William J. Jagust, Keith A. Johnson, Chester A. Mathis, Michael J. Pontecorvo, Christopher C. Rowe, Mark A. Mintun, et al. The centiloid project: Standardizing quantitative amyloid plaque estimation by PET. Alzheimer’s & Dementia, 11(1):1–15.e4, 2015. doi: 10.1016/j.jalz.2014.07.003. Michael Navitsky, Abhinay D. Joshi, Ian Kennedy, William E. Klunk, Christopher C. Rowe, Dean F. Wong, Howard Aizenstein, Michael D. Devous, Mark A. Mintun, Michael J. Pontecorvo, et al. Standardization of amyloid quantitation with florbetapir standardized uptake value ratios to the centiloid scale. Alzheimer’s & Dementia, 14(12):1565–1571, 2018. doi: 10.1016/j.jalz.2018.06. 1353. Jerzy Neyman. On the two different aspects of the representative method: The method of stratified sampling and the method of purposive selection. Journal of the Royal Statistical Society, 97(4): 558–625, 1934. doi: 10.2307/2342192. 10 Sebastian Palmqvist, Pontus Tideman, Niklas Mattsson-Carlgren, Suzanne E. Schindler, Ruben Smith, Rik Ossenkoppele, Olof Strandberg, Shorena Janelidze, Erik Stomrud, Henrik Zetterberg, Kaj Blennow, Randall J. Bateman, Oskar Hansson, et al. Blood biomarkers to detect alzheimer disease in primary care and secondary care. JAMA, 332(15):1245–1257, 2024. doi: 10.1001/jama. 2024.13855. Sebastian Palmqvist et al. Alzheimer’s association clinical practice guideline on the use of blood- based biomarkers in the diagnostic workup of suspected alzheimer’s disease within specialized care settings. Alzheimer’s & Dementia, 2025. doi: 10.1002/alz.70535. Hong Qian. Open-system nonequilibrium steady state: Statistical thermodynamics, fluctuations, and chemical oscillations. The Journal of Physical Chemistry B, 110(31):15063–15074, 2006. doi: 10.1021/jp061858z. Hong Qian. Phosphorylation energy hypothesis: Open chemical systems and their biological func- tions. Annual Review of Physical Chemistry, 58:113–142, 2007. doi: 10.1146/annurev.physchem. 58.032806.104550. James M. Robins, Andrea Rotnitzky, and Lue Ping Zhao. Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association, 89(427): 846–866, 1994. doi: 10.1080/01621459.1994.10476818. Carl-Erik Särndal, Bengt Swensson, and Jan Wretman. Model Assisted Survey Sampling. Springer, 1992. Reisa A. Sperling, Elizabeth C. Mormino, Aaron P. Schultz, Rebecca A. Betensky, Kathryn V. Papp, Rebecca E. Amariglio, Bernard J. Hanseeuw, Rachel Buckley, Jasmeer Chhatwal, Trey Hedden, Gad A. Marshall, Keith A. Johnson, et al. Association of factors with elevated amyloid burden in clinically normal older individuals. JAMA Neurology, 77(6):735–745, 2020. doi: 10. 1001/jamaneurol.2020.0387. Reisa A. Sperling, Michael C. Donohue, Rema Raman, Michael S. Rafii, Keith A. Johnson, Colin L. Masters, Christopher H. van Dyck, Takeshi Iwatsubo, Gad A. Marshall, Paul S. Aisen, et al. Trial of solanezumab in preclinical alzheimer’s disease. New England Journal of Medicine, 389 (12):1096–1107, 2023. doi: 10.1056/NEJMoa2305032. U.S. Food and Drug Administration. KISUNLA (donanemab-azbt) prescribing information. FDA label, 2024. URL https://w.accessdata.fda.gov/drugsatfda_docs/label/2024/ 761248s000lbl.pdf. Accessed 2026-08-23. U.S. Food and Drug Administration. LEQEMBI (lecanemab-irmb) prescribing information. FDA label, 2026. URL https://w.accessdata.fda.gov/drugsatfda_docs/label/2026/ 761269s011lbl.pdf. Accessed 2026-08-23. 11