Paper deep dive
Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling
Clemens SchÀchter, Astrid Pechmann, Janbernd Kirschner, Jan Hasenauer, Harald Binder
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/23/2026, 2:17:14 AM
Summary
The paper proposes using Large Language Models (LLMs) as synthetic clinical experts to supervise variational autoencoders (VAEs) for modeling longitudinal rare-disease data. Specifically, LLMs are queried to judge clinical categories (e.g., SMA types) from patient observations. These judgments are distilled into a differentiable surrogate model, which adds a Jensen-Shannon divergence penalty to the VAE loss function. This ensures that reconstructed patient profiles maintain clinical consistency (faithfulness) with the original observations. The method was applied to spinal muscular atrophy (SMA) motor-function data, reducing label disagreement from 11% to 7% and improving the prediction of motor function milestones compared to unsupervised baselines.
Entities (8)
Relation Signals (8)
Large Language Models â usedas â Synthetic Clinical Expert
confidence 98% · we use large language models (LLMs) as synthetic clinical experts to supervise a variational-autoencoder-based approach
Proposed Method â appliedto â Spinal Muscular Atrophy
confidence 97% · In an application to longitudinal motor-function assessments from children with spinal muscular atrophy
Large Language Models â supervises â Variational Autoencoder
confidence 95% · LLMs are queried offline on textual descriptions of patient observations to obtain judgments... To improve the variational autoencoder fit, we train a differentiable surrogate model on these judgments
Surrogate Model â approximates â Large Language Models
confidence 94% · train an artificial neural network on the LLM judgments to obtain a differentiable surrogate of the synthetic expert
Surrogate Model â augments â Loss Function
confidence 93% · augment the loss function to encourage reconstructions that preserve the clinical-label distribution
Jensen-Shannon Divergence â usedin â Loss Function
confidence 92% · include this term as a penalty in the loss function of R Ï ... L SE R (Ï; y, Ë y, s) =L R (Ï; y, Ë y, s) + λ SE JS (Ï, Ë Ï)
Proposed Method â improves â Motor Function Milestones Prediction
confidence 91% · informing the latent representation by the synthetic expert improved prediction of motor function milestones
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, elicitation of expert knowledge and formalization for model fitting is challenging, in particular due to limited time of clinical experts. To nevertheless make domain knowledge accessible during model fitting, we use large language models (LLMs) as synthetic clinical experts to supervise a variational-autoencoder-based approach that learns low-dimensional latent summaries of visit-level observations. Specifically, LLMs are queried offline on textual descriptions of patient observations to obtain judgments, e.g., the suspected clinical category. To improve the variational autoencoder fit, we train a differentiable surrogate model on these judgments and augment the loss function to encourage reconstructions that preserve the clinical-label distribution of their corresponding input profile. In an application to longitudinal motor-function assessments from children with spinal muscular atrophy, we map visit-level clinical profiles to low-dimensional representations that are linked by a multivariate mixed-effects model. The synthetic expert loss discourages reconstructions that remain numerically close in data space but alter the clinical interpretation of the reconstructed motor function profile, such as by crossing a disease-type boundary. We thus reduced disagreement between original and reconstructed SMA type labels from about 11 to 7 percent. Furthermore, informing the latent representation by the synthetic expert improved prediction of motor function milestones compared with unsupervised latent representations and a data-level baseline. These results suggest that incorporating LLMs into model fitting can make clinical knowledge available to representation learning and improve clinical faithfulness for longitudinal rare-disease data.
Tags
Links
- Source: https://arxiv.org/abs/2608.16507v1
- Canonical: https://arxiv.org/abs/2608.16507v1
Trouble viewing inline? Open PDF directly â
Full Text
74,594 characters extracted from source content.
Expand or collapse full text
Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling Clemens SchĂ€chter 1,2,* , Astrid Pechmann 3 , Janbernd Kirschner 3 , Jan Hasenauer 4,5 , and Harald Binder 1,2,6 1 Institute of Medical Biometry and Statistics (IMBI), Faculty of Medicine and Medical Center â University of Freiburg, Freiburg, Germany 2 Freiburg Center for Data Analysis, Modeling and AI â University of Freiburg, Freiburg, Germany 3 Department of Neuropediatrics and Muscle Disorders, Faculty of Medicine and Medical Center â University of Freiburg, Freiburg, Germany 4 Bonn Center for Mathematical Life Sciences â University of Bonn, Germany 5 Life and Medical Sciences (LIMES) Institute â University of Bonn, Germany 6 CIBSS, Centre for Integrative Biological Signalling Studies â University of Freiburg, Freiburg, Germany * Corresponding author 1 arXiv:2608.16507v1 [cs.AI] 17 Aug 2026 Summary Due to the limited amount of information, modeling longitudinal rare-disease data can particularly benefit from integrating clinical knowledge. Yet, elicitation of expert knowledge and formalization for model fitting is challenging, in particu- lar due to limited time of clinical experts. To nevertheless make domain knowl- edge accessible during model fitting, we use large language models (LLMs) as synthetic clinical experts to supervise a variational-autoencoder-based approach that learns low-dimensional latent summaries of visit-level observations. Specif- ically, LLMs are queried offline on textual descriptions of patient observations to obtain judgments, e.g., the suspected clinical category. To improve the variational autoencoder fit, we train a differentiable surrogate model on these LLM-derived judgments and augment the loss function to encourage reconstructions that pre- serve the clinical-label distribution of their corresponding input profile. In an ap- plication to real-world longitudinal motor-function assessments from children with spinal muscular atrophy in the SMArtCARE clinical registry, we map visit-level clin- ical profiles to low-dimensional representations that are linked by a multivariate mixed-effects model. The synthetic expert loss discourages reconstructions that remain numerically close in data space but alter the clinical interpretation of the re- constructed motor function profile, such as by crossing a disease-type boundary. We thus reduced disagreement between original and reconstructed SMA type la- bels from about 11 to 7 percent. Furthermore, informing the latent representation by the synthetic expert improved prediction of motor function milestones as exter- nal endpoints compared with unsupervised latent representations and a data-level baseline. These results suggest that incorporating LLMs into model fitting can make clinical knowledge available to representation learning and improve clinical faithfulness for longitudinal rare-disease data. Keywords: Generative models; Large language models; Longitudinal modeling; Spinal muscular atrophy; Synthetic clinical experts; Variational autoencoders 1 Introduction To incorporate clinical knowledge into statistical modeling, clinical experts need to devote time to interacting with modelers, and the latter must find ways to formalize the expert input into model structure or model fitting approaches. Clinical interpre- tation in rare-disease patient cohorts is often shaped by experience and contextual information to compensate for scarce data. Therefore, knowledge about distinc- 1 tions between trajectories of patients that are meaningful in clinical practice may be hard to formalize as explicit formulas, priors, or constraints (Daee et al. 2017; Mikkola et al. 2023). Furthermore, the required sustained exchange between clini- cal and statistical experts makes the formalization time-consuming and vulnerable to errors arising from miscommunication across disciplines. As a result, statisti- cal models for clinical data are often just optimized for objectives that are easier to formalize, such as the likelihood of data given a set of model parameters, pre- diction performance concerning some clinical endpoint, or reconstruction loss in representation learning. While these objectives are the established workhorses of statistical modeling, they can fail to preserve clinically relevant structure when that structure is not explicitly encoded in the optimization objective (Kelly et al. 2019; Oakden-Rayner et al. 2019; Vickers and Elkin 2006). A novel approach to tackle this challenge could be offered by large language models (LLMs), which incorporate extensive biomedical knowledge from their train- ing on clinical literature and can retrieve clinical information in natural language (Singhal, Azizi, et al. 2023; Thirunavukarasu et al. 2023). Specifically, LLMs oper- ationalize biomedical knowledge in text-continuation tasks, which has already been used successfully in medical explanation and question-answering tasks (Singhal, Tu, Gottweis, et al. 2025; Thirunavukarasu et al. 2023) and in clinical text sum- marization (Veen et al. 2024). Incorporating question-answering tasks into model fitting provides an opportunity to make clinical knowledge available to statistical modeling in an automated, scalable manner. We specifically consider LLMs as synthetic clinical experts for generative models, namely conditional variational au- toencoders (cVAEs) (Kingma and Welling 2013; Sohn, Lee, and Yan 2015), that provide dimension reduction for each time point where a patient has been ob- served. Here, LLM-derived judgments are used to encourage the cVAE to pre- serve clinical interpretations during reconstruction. This is motivated by an application to a longitudinal dataset of children with spinal muscular atrophy (SMA), a rare neuromuscular disorder characterized by progressive muscle weakness and impaired motor-function development (Mercuri, Sumner, et al. 2022; Schorling, Pechmann, and Kirschner 2020). Clinically, SMA severity is often summarized via three broad SMA type categories, which group patients by expected disease severity and motor milestone attainment (Calucho et al. 2018; Mercuri, Finkel, et al. 2018; Varone, Esposito, and Bitetti 2025). These categories support clinical description, define patient groups in studies, and inform clinical management and treatment planning (Mercuri, Finkel, et al. 2018; Schor- ling, Pechmann, and Kirschner 2020; Varone, Esposito, and Bitetti 2025). In our application, we focus on motor function of patients, which is measured using a 2 specialized assessment (Bishop, Montes, and Finkel 2018; Haataja et al. 1999) that records multiple ordinal test items covering a wide range of abilities such as head control, sitting, crawling, standing, and walking (Bishop, Montes, and Finkel 2018; Haataja et al. 1999). Since motor function profiles are interpreted relative to the childâs age and expected attainment of developmental milestones, superfi- cially similar profiles may receive different SMA type labels from a clinical expert (Bishop, Montes, and Finkel 2018; Calucho et al. 2018), such that no straightfor- ward formalization is available, but an LLM might nevertheless be able to mimic expert assessment. To model the development of motor function, we consider a VAE-based repre- sentation approach in which a cVAE learns visit-level latent variables and a mul- tivariate mixed-effects model links these variables across observation times (Ong et al. 2024; SchĂ€chter et al. 2025). Dimension reduction by a cVAE helps to keep the number of parameters moderate in the face of a large number of motor func- tion test items, which is particularly relevant in settings with small sample sizes, such as in a rare disease like SMA. However, a subsequent reconstruction from a latent variable may be numerically close to the observed profile while neverthe- less crossing a clinically meaningful decision boundary. Conversely, a numerically larger reconstruction error may be less important when the reconstructed profile retains the same functional interpretation. To incorporate such potential differences in clinical meaningfulness into model fitting, LLMs can provide assessments of the original and reconstructed observa- tions, and any discrepancies can inform the cVAE representation. For assessment by LLMs, we formulate pairwise question-answering tasks in which the LLM must assign a patientâs motor function profile to one of two candidate SMA type labels, separately for the original and reconstructed data. Performing these pairwise tasks over all combinations of SMA type labels allows us to identify inconsistencies in LLM answers before using them to guide the cVAE. To actually shape the cVAE representation, we need a differentiable criterion for augmenting the loss function. Therefore, we train an artificial neural network on the LLM judgments to obtain a differentiable surrogate of the synthetic expert that can provide gradients during VAE training. Our approach is related to work that uses LLM outputs not only as final pre- dictions, but as supervision for subsequent learning procedures. In LLM-based weak supervision, prompted LLMs can be used to define labeling functions or generate weak labels that are then used to train downstream classifiers (Hsu and Roberts 2025; Smith et al. 2022). Related distillation work uses LLM-generated labels and rationales as supervision for smaller task-specific models (Hsieh et al. 3 2023). This is particularly relevant in clinical settings, where expert annotation is expensive and many judgments depend on contextual interpretation. However, such approaches usually use such labels to produce training data for classifiers or downstream predictors, rather than to steer the model fitting process of a sta- tistical or generative model. Here, an additional challenge arises, because LLM judgments can be expensive, stochastic, and non-differentiable with respect to the parameters of the model that should be improved. Our use of a surrogate model for directly influencing model fitting follows a broader strategy in which complex, non- differentiable supervision is first converted into an auxiliary model before being used for optimization. Examples include knowledge distillation, where a student model is trained to approximate information provided by a larger teacher model (Hinton, Vinyals, and Dean 2015; Hsieh et al. 2023), and preference-based rein- forcement learning, where comparison feedback is used to learn a reward model for policy optimization (Christiano et al. 2017; Ouyang et al. 2022). This work extends LLM-based supervision beyond downstream label genera- tion by using LLM-derived judgments directly in the fitting objective of a gener- ative representation model to discourage clinically inconsistent reconstructions. We investigate whether this supervision can improve the clinical faithfulness and milestone-prediction utility of latent representations learned from irregular longi- tudinal real-world SMA registry data. Section 2 formalizes our framework by in- troducing LLMs acting as synthetic clinical experts and then describing how their judgments are audited for internal consistency. These judgments are subsequently distilled into a differentiable surrogate model, and incorporated into VAE train- ing to penalize reconstructions that remain numerically close to the input but al- ter the clinical interpretation of the profile. Section 3 evaluates the approach in longitudinal SMA motor-function data, where clinical categories are contextual, age-dependent, and difficult to encode as explicit rules, and examines whether synthetic-expert supervision reduces clinically inconsistent reconstructions and improves the prognostic usefulness of the resulting latent representations for motor milestone prediction. Section 4 considers methodological limitations and potential extensions. 4 2 Methods 2.1 Synthetic-expert label generation When clinical experts assess a clinical observation y â Y with context s â S and assign it to a clinical label c h â C = c 1 ,...,c H , they can follow fixed rules but may also rely on domain knowledge and experience to form their judgment. LLMs acting as synthetic clinical experts provide the possibility to retrieve such assessments based on domain knowledge in an automated manner and with short turnaround time, when compared to real clinical experts. This opens a new way for incorporating knowledge in settings where fitted models are generative, i.e. provide means for sampling data from them: Instead of mathematically formalizing knowledge, e.g. as priors for parameter estimation, the LLM can be used to judge the sampled data in terms of clinical interpretability. Formally, LLMs work by segmenting input text into a sequence of tokens, which then are mapped to a high-dimensional space by an embedding model, resulting in a sequence of vectors v â ,â = 1,...,L. These serve as the input for a transformer neural network f LLM (·), whose parameters were trained on a large corpus of text. While an LLM can generate arbitrary words as output, we constrain it (via corre- sponding instructions) to a forced-assignment task to one out of a pair of clinical labels. This question format is related to paired-comparison approaches in psycho- metrics and to relative-similarity queries, where judgments are elicited by asking which of two alternatives better matches a reference case rather than by assigning an absolute category rating (Tamuz et al. 2011; Thurstone 1927). Let H denote the number of clinical categories in C. For each ordered pair- wise combination of labels m = 1,...,M = H(H â 1), let (a m ,b m ) with a m ,b m â 1,...H,a m Ìž= b m denote the two label indices with corresponding candidate la- bels c a m ,c b m âC. For a rendered clinical profile and a specified pair of candidate labels, the LLM- induced pairwise judgment can then be written as c (m) LLM = f LLM v instr,m 1 ,..., v instr,m L 1 , v data 1 ,..., v data L 2 , c (m) LLM âc a m ,c b m , m = 1,...,M. (1) where L 1 of the L input tokens correspond to the instructions (including specifi- cation of the respective pair m), and L 2 tokens correspond to the data y to be evaluated, with L = L 1 + L 2 . 5 For obtaining v instr,m â , â = 1,...,L 1 , and v data â , â = 1,...,L 2 , from the embedding model, a decision is needed about what text to feed into the latter. The instruction text can be augmented with complementary visit information s that is relevant to form a clinical judgment, such as patient age or other covariates. To render the data y as text, each clinical variable can be paired with its observed value and a textual description of the measured quantity. In some settings, numerical scores are ordinal codes whose meaning is defined by the clinical description rather than the numerical magnitude alone. For example, in the SMA application, HINE-2 motor-function scores correspond to specific motor function abilities of the child. In such cases, it may be preferable to provide the textual item descriptions instead of the raw numeric scores. One advantage of the forced-assignment design is that the consistency of the LLM responses c (m) LLM ,m = 1,...,M can be evaluated directly. First, stability un- der label-order reversal can be assessed by comparing a pair of candidate label indices (a m ,b m ) with its reversed counterpart (b m ,a m ). A response is considered stable if the same clinical label is selected in both directions. Furthermore, when the clinical labels can be arranged in an ordinal order, the full set of pairwise answers can be checked for ordinal consistency. Let c 1 âș c 2 âș · âș c H denote the clinical label order, for example SMA type 1 âș SMA type 2 âș SMA type 3âș presymptomatic motor development in our SMA application. For a given profile, the pairwise answers are called ordinal consistent if there exists a single label index k â1,...,H such that every pairwise choice is compatible with c k being the closest label on this ordered scale. Formally, for each unordered com- parison m with a m < b m , choosing c a m over c b m imposes k †(a m + b m )/2, whereas choosing c b m over c a m imposes k â„ (a m + b m )/2. If no such k exists, the answer pattern is classified as internally inconsistent, whereas consistent answers are ag- gregated into a final synthetic-expert label c SE = c k â C. This ordinal-consistency check can be applied separately to the answers obtained under each candidate- label order. 2.2 Synthetic-expert supervision We next describe how the synthetic-expert judgments enter the model-fitting ob- jective. For this, we consider a model R Ï with parameters Ï that, for an observed clinical profile yâY and context sâS, produces a reconstruction Ë y = R Ï (y, s)âY, with Ë yâ y. 6 This notation includes deterministic reconstruction maps as well as probabilistic generative models, in which Ë y denotes the decoder mean, predicted item proba- bilities, or a differentiable reconstruction summary used in the loss function. We denote the synthetic-expert label of the observed profile by c SE â C and the corresponding label of the reconstruction by Ëc SE â C. We want to use the synthetic expert to steer the model R Ï toward reconstructions for which Ëc SE aligns with c SE . Since the parameters Ï of the generative model R Ï are usually obtained by minimizing its loss function L R (Ï; y, Ë y, s) via gradient-based methods, we re- quire a differentiable approximation to the synthetic expertâs class-score function. Therefore, we train a surrogate neural network S Ï that approximates the synthetic- expert label given a sample y and contextual information s and maps this input to a probability distribution over the clinical labels: S Ï (y, s) = (Ï 1 ,...,Ï H ) =Ï â â Hâ1 , Ï h â Pr(c SE = c h | y, s), h = 1,...,H, (2) where Ï h approximates the probability that the LLM-derived final clinical label for profile y and contextual information s is c h . A sufficiently large training set for the surrogate can be generated by querying the LLM on real observations, on synthetic data sampled from a prefitted generative model, or on a combination of both. To measure inconsistency between the surrogate distribution for the observed profileÏ = S Ï (y, s) â â Hâ1 and the corresponding distribution for the reconstruc- tion Ë Ï = S Ï ( Ë y, s)â â Hâ1 , we calculate the Jensen-Shannon divergence, JS (Ï, Ë Ï) = 1 2 KL (Ï || m) + 1 2 KL ( Ë Ï || m),m = 1 2 (Ï + Ë Ï).(3) where KL(·||·) is the Kullback-Leibler divergence. Finally, we include this term as a penalty in the loss function of R Ï to disincen- tivize reconstructed profiles for which the reconstruction label Ëc SE differs from the original label c SE , L SE R (Ï; y, Ë y, s) =L R (Ï; y, Ë y, s) + λ SE JS (Ï, Ë Ï).(4) The scaling factor λ SE weights the influence of the supervision term relative to the rest of the loss function. During optimization of R Ï , the surrogate parameters Ï are held fixed, although gradients are propagated through the surrogate with respect to its reconstructed-profile input. 7 2.3 Linking visit-level latent representations with a multivari- ate mixed-effects model Having defined the synthetic-expert supervision signal, we next incorporate it into a representation model for longitudinal rare-disease data. For this, let i â I denote an individual for which longitudinal measurements y i,t â R n are documented at observation times T i = t i,1 ,...,t i,m i , with t i,1 < · < t i,m i . In rare-disease studies, the number of individuals may be small relative to the dimensionality and complexity of the visit-level measurements Consequently, directly modeling high-dimensional visit-level observations with a statistical model such as a multivariate longitudinal mixed model can be infeasible. We therefore use a conditional variational autoencoder (cVAE) (Kingma and Welling 2013; Sohn, Lee, and Yan 2015) to represent each observation through a lower-dimensional la- tent variable z i,t â R d , with d < n, such that statistical modeling becomes feasible. For this latent representation, we use the standard Gaussian prior p(z i,t ) =N d (0, I d ). Because the synthetic-expert judgment depends on contextual information such as age, both encoding and decoding are conditioned on s i,t . The encoder defines the approximate posterior q Ï (z i,t | y i,t , s i,t ) =N d ÎŒ i,t , diag(Ï 2 i,t ) , whereÎŒ i,t andÏ i,t are neural-network outputs. To link longitudinal latent representations Z i = (z i,t ) †tâT i â R m i Ăd , we follow our previous work (SchĂ€chter et al. 2025) and model them by multivariate mixed- effects regression, Z i = X i B + T i U i + E i .(5) Here X i â R m i Ăp is the fixed-effect design matrix, B â R pĂd contains population- level effects, T i â R m i Ăq is the random-effect design matrix, U i â R qĂd contains individual-specific deviations, and E i â R m i Ăd is residual error with vec(U i )âŒN qd (0, Ί),vec(E i )âŒN m i d (0, ÎŁâ I m i ),(6) where Ίâ R qdĂqd and ÎŁâ R dĂd are covariance matrices. For example, in the SMA application, X i may include visit age, genetic markers, treatment indicators and demographic baseline patient characteristics. Similarly, T i may include a patient- 8 specific intercept and slope. Given estimates ( Ë B, Ë ÎŠ, Ë ÎŁ) and a mixed model prediction Ë z i,t for z i,t , the decoder defines the conditional reconstruction distribution p Ξ (y i,t | Ë z i,t , s i,t ), with the distributional form chosen to match the clinical scale of each component of y i,t . The approach is fitted using an estimation scheme in which the encoder and decoder parameters (Ï,Ξ) are updated alternately with the latent mixed-effects model parameters (B, Ί, ÎŁ). First, the encoder and decoder parameters are up- dated while the latent mixed-effects model parameters are frozen. While updating the encoder and decoder, we use the surrogate network to encourage consistency between observed and reconstructed profiles. Specifi- cally, we optimize a ÎČ-VAE-style objective (Higgins et al. 2017) augmented by two penalties: one aligning the encoder representation with the latent mixed-model prediction, and one aligning the surrogate distributions of the observed and recon- structed profiles: L SE cVAEâM Ï,Ξ = X iâM 1 |T i | X tâT i â E z i,t âŒq Ï (z i,t |y i,t ,s i,t ) [logp Ξ (y i,t | Ë z i,t , s i,t )] + ÎČKL [q Ï (z i,t | y i,t , s i,t )|| p(z i,t )] + Î·â„ Ë z i,t â z i,t â„ 2 2 + λ SE JS (S Ï (y i,t , s i,t ),S Ï ( Ë y i,t , s i,t )) ! , (7) where M â I is a mini-batch. We define the reconstruction Ë y i,t as the conditional mean of the decoder distribution p Ξ (·| Ë z i,t , s i,t ), rather than as a sample drawn from this distribution. This keeps the surrogate-based supervision term differentiable during model fitting. After some iterations, the encoder and decoder parameters are frozen and latent representations are obtained for all observations. Subse- quently, the latent mixed-model parameters (B, Ί, ÎŁ) are estimated via maximum likelihood. These two optimization steps are alternated for a prespecified number of epochs or until a stated convergence criterion is met. After the training process has concluded, the latent variables z i,t provide low- dimensional summaries of the visit-level observations y i,t . Using them in down- stream tasks, such as time-to-event modeling, provides a way to assess whether the learned variables retain information relevant for predicting future events. Specif- ically, for an event time T i and censoring indicator ÎŽ i , the observations of patient i can be represented as time-varying intervals on the observation-time scale. On 9 an interval (t i,j ,t i,j+1 ], a Cox model on the latent representation can be written as λ i (t| z i,t i,j ) = λ 0 (t) exp ÎČ â€ z i,t i,j , tâ (t i,j ,t i,j+1 ].(8) Here, λ 0 (t) denotes the baseline hazard andÎČ quantifies the association between the learned latent representation and the event hazard. In the SMA application, this Cox model is used to evaluate whether the latent representation contains in- formation about future attainment of motor-function milestones, such as sitting, standing, and walking. Figure 1 provides a graphical summary of the proposed workflow. It highlights the separation between offline LLM-based label generation and the differentiable training loop, where the surrogate provides the path through which synthetic-expert feedback enters the cVAE loss. 3 Results 3.1 Application to spinal muscular atrophy data We evaluated synthetic-expert supervision using longitudinal motor-function as- sessments from the SMArtCARE registry (Pechmann et al. 2019), a clinical reg- istry for patients with spinal muscular atrophy (SMA). In the SMArtCARE registry, disease severity is monitored through motor-function assessments at repeated follow-up visits using instruments such as the HINE-2 motor-function test (Bishop, Montes, and Finkel 2018; Haataja et al. 1999). During these assessments, pa- tients perform several motor-function tasks that are graded by physicians on or- dinal item scales. These test results should not be interpreted in isolation: their clinical meaning depends on the childâs age and developmental context and similar HINE-2 profiles can therefore carry different clinical implications. To model disease progression for these visit-level profiles, we used the cVAE- based representation approach introduced in Section 2.3. The cVAE compresses each HINE-2 profile, conditioned on the corresponding age value, into a low- dimensional latent representation and reconstructs the original item profile from this representation, while the longitudinal multivariate mixed-effects model links the latent variables across observation times. For the synthetic-expert task, we used the three conventional symptomatic SMA type categories, SMA types 1â3, which reflect typical age at onset and ex- pected motor-function progression (Mercuri, Finkel, et al. 2018). SMA type 1 typ- ically begins in early infancy and is associated with severe muscle weakness and 10 failure to achieve independent sitting. SMA type 2 usually describes patients who achieve sitting but not independent walking. SMA type 3 has symptom onset after infancy, and patients usually acquire independent walking for at least part of the disease course. We additionally included presymptomatic motor development as a fourth category for observations from patients who had been diagnosed with SMA but did not yet show clear clinical signs. We used synthetic-expert supervision to disincentivize decoder reconstructions that reproduce the HINE-2 profiles accurately while changing the clinical reading of the profile in terms of these categories, for example by shifting an age-appropriate pattern toward delayed motor development or by moving a profile slightly across an SMA type boundary. Specific implementation details for preprocessing, synthetic-expert generation, surrogate training, and model fitting are provided in Appendix A. 3.2 Obtaining a differentiable synthetic-expert surrogate on SMA data We first assessed whether LLM-derived judgments on SMA motor-function pro- files could be turned into a reliable differentiable supervision signal. For this pur- pose, we queried LLMs from the Qwen-3 (A. Yang et al. 2025), GPT-OSS (OpenAI 2025), and Gemma-4 (Google DeepMind 2026) model families to classify textual descriptions of HINE-2 profiles paired with the patientâs age. Each profile was eval- uated through pairwise comparisons between the four labels, SMA types 1-3 and presymptomatic motor development. The pairwise responses were then mapped to a final SMA type label by considering the ordinal ordering of the SMA type cat- egories. We used three criteria to evaluate consistency of the LLM answers. First, we used the pairwise comparison design to investigate ordinal consistency by count- ing contradictory patterns. For example, when an LLM selects "SMA type 1" over "SMA type 2", but also "SMA type 3" over "SMA type 2" for the same profile, the resulting pattern is classified as inconsistent because it violates the ordinal struc- ture of the SMA type labels. Second, we evaluated robustness to prompt wording by reversing the order in which the candidate labels were presented to the LLM. Finally, we queried the LLM to return an uncertainty score for each answer on a low, medium, or high scale and examined whether uncertainty was elevated for samples in which the LLM failed to recover the same label after label-order rever- sal. Table 1 summarizes the internal-consistency metrics for the evaluated LLM 11 raters and for a baseline that independently selects either candidate with prob- ability p = 0.5. Small models, such as Qwen-3-0.6B, showed output patterns that were close to random baseline, indicating that their pairwise label choices con- tained little stable clinical structure. With increasing model size, the consistency of the synthetic-expert judgments improved noticeably. Gemma-4-31B had the highest overall internal-consistency and order-stability metrics and yielded consis- tent final labels for around 89.9% of samples and agreement under label-order reversal for 97.8% of pairwise comparisons and 92.7% of final SMA type labels. Notably, the stronger models also assigned higher self-reported uncertainty to an- swers that were not stable under label-order reversal. This effect was not apparent for smaller models such as Qwen-3-0.6B and Qwen-3-1.7B, but was 9.6 times higher for Gemma-4-31B among order-unstable comparisons than among stable comparisons. Subsequently, the surrogate model was trained to reproduce the final class label from age and the thermometer-encoded HINE-2 profile with two heads to reproduce both label-order choices. Profiles with contradictory pairwise decision patterns were excluded from the training set allowing the surrogate to interpolate for samples with inconsistent pairwise answers. For the largest model in each fam- ily, surrogate accuracy exceeded 90% in held-out folds in fivefold cross-validation (Table 1). The surrogate provides a fast, differentiable approximation to the input-output mapping induced by the selected LLM, prompt, decoding procedure, and aggre- gation rule. To visualize this, we evaluated the decision boundaries for the sit, stand and walk motor-function items. Here we grouped observations into motor function profiles that achieved independent sitting but not standing, independent standing but not walking, and walking. We varied the age values and averaged the predicted class probabilities over all samples of the dataset. We show these results in Figure 2. We observed that the same motor-function profile receives different class probabilities depending on age, showing that the surrogate cap- tures the developmental context used by the synthetic expert models. Consistent with the clinical expectation, motor function profiles of patients that have sitting as highest motor function are eventually classified as SMA type 2 after being labeled presymptomatic at a young age. Patients who can stand but not walk are classified as either SMA type 2 or SMA type 3, whereas patients who can walk are predom- inantly classified as presymptomatic. The latter group typically has no recorded motor deficits on the HINE-2 because these patients reach the upper limit of the scale. We also use the synthetic expert to make latent variables interpretable in terms of SMA type categories, by decoding latent variables and classifying the 12 generated samples by the surrogate network. This allows us to assess latent de- cision boundaries. We show these results in Figure 2. 3.3 Consequences for consistency and reconstruction loss We next evaluated how synthetic-expert supervision affected the reconstruction behavior of the cVAE-based model. For this analysis, we used surrogate networks derived from the most consistent LLMs within each evaluated model family. Specifically, we tracked the Jensen-Shannon divergence between the categor- ical surrogate judgment distributions and the proportion of profiles for which the observed and reconstructed profiles received different expected SMA type labels from the surrogate model. To assess whether improvements in SMA-type label consistency came at the cost of numerical reconstruction quality, we additionally recorded the average HINE-2 sum score reconstruction error. Finally, we also evaluated the reconstruction error for profiles with initial label disagreement un- der the unsupervised baseline. Analyses were repeated over 10 random seeds and across different choices of supervision weights λ SE â 0, 10, 25, 50, 100. The resulting supervision sweep is summarized in Table 2. Across the LLM model families, increasing the supervision weight reduced the discrepancy between observed profiles and their reconstructions in the surrogate judgment space. We observed this pattern both for the distributional comparison of class probabilities via JS divergence and for the total error rate in expected SMA type labels. Therefore, the supervision signal showed a consistent shift toward re- constructions that retained the same clinical interpretation by the surrogate model as the corresponding input profiles. This gain in consistency with the LLM-derived functional categorization was not necessarily accompanied by a substantial trade-off in reconstruction accuracy. Across the tested supervision weights, the average HINE-2 reconstruction error remained close to the unsupervised baseline, with small deterioration in recon- struction quality for larger supervision weights. However, profiles whose unsu- pervised reconstructions crossed an SMA type boundary under the unsupervised cVAE baseline saw a noticeable increase in reconstruction quality. The supervi- sion term seemed to shift the model capacity toward samples whose reconstruc- tions cross a clinical label boundary. As a consequence, samples without label disagreement received less emphasis in the fitted objective, leading to a small deterioration in their numerical reconstruction quality. 13 3.4 Milestone prediction with Cox regression Finally, we used a Cox regression to assess whether the supervised latent rep- resentations carried prognostic information for future functional milestones. The SMArtCARE dataset tracks sitting, standing, and walking milestones, which are related to the corresponding HINE-2 items but not identical to them. For each milestone, visits were represented as time-varying Cox intervals on the age scale. The fitted models predicted milestone attainment within one year after each held- out landmark visit. Performance was compared with baselines based on the HINE-2 sum score, the HINE-2 sum score plus fixed clinical covariates, and the latent representation of the cVAE base model. Supervised latent representations were evaluated across synthetic-expert model families and supervision weights (Table 3). We evaluated model performance using three metrics. First, we used the IPCW Brier score to quantify the pointwise accuracy of predicted one-year milestone risk. Second, we used the time-dependent AUC to assess discrimination between patients who attained the milestone within one year and those who did not. Third, we evaluated calibration by computing the weighted mean absolute difference between predicted and observed one-year risk across calibration bins. All metrics were averaged over 10 random seeds and computed with patient-grouped five-fold cross-validation. We observed that the Cox models fitted on the latent representation of the cVAE base model outperformed the data-level baselines on the reported metrics. The supervised variants generally preserved this prognostic information and further im- proved performance relative to the unsupervised latent baseline. This pattern is consistent with the role of the synthetic-expert signal: because SMA type cate- gories are partly defined by expected milestone attainment, a supervision term that encourages clinically consistent SMA type interpretation may also encour- age the latent representation to retain information relevant for milestone predic- tion. However, the supervision weight should be chosen carefully; for larger values like λ SE = 100 we observed a decrease in performance on some metrics, whereas moderate weighting like λ SE = 25 appeared to provide a balance between supervi- sion signal and the original cVAE loss terms. 4 Discussion In this work, we proposed a framework for using LLM-derived clinical judgments as a supervision signal for longitudinal rare-disease modeling. We were motivated by the difficulty of incorporating clinical knowledge into model fitting, particularly when 14 relevant distinctions are well understood by experts and described in the medical literature, but are difficult to translate into explicit formulas, priors, or constraints. To address this, we proposed an approach that uses an LLM as a synthetic clin- ical expert to assign clinical labels to textual descriptions of observations. These labels are then used to train a differentiable surrogate, which provides a super- vision signal for steering the training of a generative representation model. Our approach targets reconstructions that remain close to their input in terms of nu- merical reconstruction error, but change the clinical interpretation of the original profile. We evaluated the approach on spinal muscular atrophy motor function data, where we extracted clinical judgments with a pairwise comparison prompt design. We found consistency of these judgments to be highly dependent on both model size and model family. Since the computational cost of querying a more capa- ble model comes only during offline label generation, the LLM quality should be prioritized before its outputs are used as part of a statistical loss. When our approach was applied to steer a cVAE-based approach we observed an improvement in consistency under the LLM-derived surrogate via a reduced disagreement between its judgment distributions of observed and reconstructed profiles. For moderate supervision weights HINE-2 reconstruction errors remained in a similar range, indicating that the additional loss component did not hinder the numerical reconstruction objective. However, the relevance of synthetic-expert supervision depends on whether the surrogate has learned clinically meaningful distinctions, which in turn depends on the quality of the LLM-derived judgments. We therefore used a Cox regression as a complementary analysis to investigate whether supervision also improves encoding of information relevant for future mo- tor milestone attainment. The milestone prediction results were supportive in this respect. Because SMA type categories are closely tied to expected attainment of sitting, standing, and walking milestones, improved consistency with respect to SMA type interpretation also encouraged the latent representation to more accu- rately encode milestone-relevant information. However, several methodological limitations follow from our approach. First, the synthetic expert is only as appropriate as the prompt, label definitions, textual ren- dering, and LLM used to generate the judgments. If systematic errors in the syn- thetic expert are distilled into the surrogate, they become part of the training objec- tive and shape the learned latent representation. For this reason, synthetic-expert outputs should be carefully audited before deploying the approach in a clinical set- ting. We used consistency checks to identify internal contradictions, sensitivity to label order, and unstable answer patterns. However, these checks cannot prove 15 that every retained label is clinically correct. Moreover, uncertainty of the synthetic expert is currently used mainly as a diagnostic property. Future extensions could therefore propagate uncertainty into the loss, for example by down-weighting un- certain labels or using softer target distributions. A second limitation is that the LLM-derived judgment is not used directly during cVAE training, but only through a learned differentiable surrogate. This is neces- sary because it makes supervision computationally feasible and provides gradi- ents, but it also introduces an approximation error. Especially in sparse or am- biguous regions of the HINE-2 profile space, the surrogate may smooth over un- certainty and fail to reproduce exact decision boundaries of the original synthetic expert. This could be mitigated by using uncertainty-aware surrogate models or re-querying the LLM for profiles near clinical decision boundaries. A third limitation concerns the supervision weight. Synthetic-expert supervision acts as a regularization term, and its effect depends on the balance between clin- ical consistency and numerical reconstruction. Low weights may be too weak to influence the learned representation, whereas high weights may overemphasize the synthetic label space and reduce reconstruction quality and even downstream prediction performance. The supervision weight should therefore be selected as an additional model hyperparameter, ideally using validation criteria that reflect both numerical reconstruction and the intended clinical use of the latent represen- tation. Taken together, our results support synthetic-expert supervision as a mecha- nism for incorporating clinical judgment into the disease progression modeling pro- cess. Our broader vision is not to replace clinical expertise, but to create a more practical interface between clinical reasoning and statistical modeling. Toward this goal, future work should strengthen clinical validation and propagate uncertainty from the synthetic expert into the loss. If these steps are successful, language- based clinical judgment and probabilistic modeling could become connected parts of a single workflow in rare-disease research. Funding This work was supported by the Deutsche Forschungsgemeinschaft (DFG, Ger- man Research Foundation) [Project-ID 499552394 â SFB 1597 Small Data]. 16 Ethics statement The use of anonymized data from the SMArtCARE registry was reviewed by the Ethics Committee of the University of Freiburg (reference no. 22-1430-S1-retro). The SMArtCARE registry authorized provision of the anonymized data for this re- search project. To protect participant confidentiality, all LLM inference was per- formed locally within secure computing environments. Data and code availability The SMArtCARE data cannot be made publicly available because of participant privacy restrictions. The analysis code will be made publicly available upon publi- cation and will be provided to reviewers upon request. Conflicts of interest Harald Binder serves as a Guest Editor of the Biostatistics special collection Sta- tistical Foundations of AI and Real-World Evidence Generation. The remaining authors declare no conflicts of interest. Acknowledgments Generative AI tools (ChatGPT, using GPT-5.4 and GPT-5.5) were used for lan- guage editing and for reviewing and debugging code written by the authors. No patient-level data were transmitted to these services. All AI-assisted text and code changes were reviewed, verified, and revised by the authors, who take full respon- sibility for the final manuscript and code. References Bishop, Kathie M., Montes, Jacqueline, and Finkel, Richard S. (2018). âMotor mile- stone assessment of infants with spinal muscular atrophy using the Hammer- smith Infant Neurological Exam-Part 2: Experience from a nusinersen clinical studyâ. In: Muscle & Nerve 57.1, p. 142â146. DOI: 10.1002/mus.25705. 17 Calucho, Maite, Bernal, Silvia, AlĂas, Laura, March, Ferran, VenceslĂĄ, Alba, RodrĂguez- Ălvarez, Francisco J., Aller, Elena, FernĂĄndez, Rosa M., Borrego, Salud, Mil- lĂĄn, JosĂ© M., Hernando, Ignacio, Baiget, Montserrat, and Tizzano, Eduardo F. (2018). âCorrelation between SMA type and SMN2 copy number revisited: An analysis of 625 unrelated Spanish patients and a compilation of 2834 reported casesâ. In: Neuromuscular Disorders 28.3, p. 208â215. DOI: 10.1016/j.nmd. 2018.01.003. Christiano, Paul F., Leike, Jan, Brown, Tom B., Martic, Miljan, Legg, Shane, and Amodei, Dario (2017). âDeep Reinforcement Learning from Human Preferencesâ. In: Advances in Neural Information Processing Systems. Vol. 30. Daee, Pedram, Peltola, Tomi, Soare, Marta, and Kaski, Samuel (2017). âKnowl- edge elicitation via sequential probabilistic inference for high-dimensional pre- dictionâ. In: Machine Learning. DOI: 10.1007/s10994-017-5651-7. Google DeepMind (2026). Gemma 4 Model Card. Google AI for Developers. Avail- able at: https://ai.google.dev/gemma/docs/core/model_card_4. Haataja, Leena, Mercuri, Eugenio, Regev, Rivka, Cowan, Frances, Rutherford, Mary, Dubowitz, Victor, and Dubowitz, Lilly (1999). âOptimality score for the neurologic examination of the infant at 12 and 18 months of ageâ. In: J Pediatr 135.2 Pt 1, p. 153â161. DOI: 10.1016/S0022-3476(99)70016-8. Higgins, Irina, Matthey, LoĂŻc, Pal, Arka, Burgess, Christopher P., Glorot, Xavier, Botvinick, Matthew M., Mohamed, Shakir, and Lerchner, Alexander (2017). âbeta-VAE: Learning Basic Visual Concepts with a Constrained Variational Frame- workâ. In: ICLR 2017. URL: https://openreview.net/forum?id=Sy2fzU9gl. Hinton, Geoffrey, Vinyals, Oriol, and Dean, Jeff (2015). âDistilling the Knowledge in a Neural Networkâ. In: NeurIPS Deep Learning and Representation Learning Workshop. Hsieh, Cheng-Yu, Li, Chun-Liang, Yeh, Chih-Kuan, Nakhost, Hootan, Fujii, Ya- suhisa, Ratner, Alexander, Krishna, Ranjay, Lee, Chen-Yu, and Pfister, Tomas (2023). âDistilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizesâ. In: arXiv preprint. DOI: 10.48550/ arXiv.2305.02301. Hsu, Enshuo and Roberts, Kirk (2025). âLeveraging Large Language Models for Knowledge-Free Weak Supervision in Clinical Natural Language Processingâ. In: Scientific Reports 15, p. 8241. DOI: 10.1038/s41598-024-68168-2. Kelly, Christopher J., Karthikesalingam, Alan, Suleyman, Mustafa, Corrado, Greg, and King, Dominic (2019). âKey challenges for delivering clinical impact with artificial intelligenceâ. In: BMC Medicine 17.1, p. 195. DOI: 10.1186/s12916- 019-1426-2. 18 Kingma, Diederik P. and Welling, Max (2013). âAuto-encoding variational bayesâ. In: arXiv preprint. DOI: 10.48550/arXiv.1312.6114. Mercuri, Eugenio, Finkel, Richard S., Muntoni, Francesco, Wirth, Brunhilde, Montes, Jacqueline, Main, Marion, Mazzone, Elena S., Vitale, Michael, Snyder, Barbara, Quijano-Roy, Susana, Bertini, Enrico, Davis, Richard H., Meyer, Ole, Simonds, Anita K., Schroth, Mary K., Graham, Robert J., Kirschner, Janbernd, Iannac- cone, Susan T., Crawford, Thomas O., Woods, Sunitha, Qian, Y., and Sejer- sen, Thomas (2018). âDiagnosis and management of spinal muscular atrophy: Part 1: Recommendations for diagnosis, rehabilitation, orthopedic and nutri- tional careâ. In: Neuromuscular Disorders 28.2, p. 103â115. DOI: 10.1016/j. nmd.2017.11.005. Mercuri, Eugenio, Sumner, Charlotte J., Muntoni, Francesco, Darras, Basil T., and Finkel, Richard S. (2022). âSpinal muscular atrophyâ. In: Nature Reviews Dis- ease Primers 8.1, p. 52. DOI: 10.1038/s41572-022-00380-8. Mikkola, Petrus, Martin, Osvaldo A., Chandramouli, Suyog, Hartmann, Marcelo, Pla, Oriol Abril, Thomas, Owen, Pesonen, Henri, Corander, Jukka, Vehtari, Aki, Kaski, Samuel, BĂŒrkner, Paul-Christian, and Klami, Arto (2023). âPrior knowl- edge elicitation: The past, present, and futureâ. In: arXiv preprint arXiv:2112.01380. DOI: 10.48550/arXiv.2112.01380. Oakden-Rayner, Luke, Dunnmon, Jared, Carneiro, Gustavo, and RĂ©, Christopher (2019). âHidden stratification causes clinically meaningful failures in machine learning for medical imagingâ. In: arXiv preprint arXiv:1909.12475. DOI: 10. 48550/arXiv.1909.12475. Ong, Priscilla, HauĂmann, Manuel, Lönnroth, Otto, and LĂ€hdesmĂ€ki, Harri (2024). âLatent mixed-effect models for high-dimensional longitudinal dataâ. In: arXiv preprint. DOI: 10.48550/arXiv.2409.11008. OpenAI (2025). âgpt-oss-120b & gpt-oss-20b Model Cardâ. In: arXiv preprint arXiv:2508.10925. Ouyang, Long, Wu, Jeff, Jiang, Xu, Almeida, Diogo, Wainwright, Carroll L., Mishkin, Pamela, Zhang, Chong, Agarwal, Sandhini, Slama, Katarina, Ray, Alex, Schul- man, John, Hilton, Jacob, Kelton, Fraser, Miller, Luke, Simens, Maddie, Askell, Amanda, Welinder, Peter, Christiano, Paul, Leike, Jan, and Lowe, Ryan (2022). âTraining Language Models to Follow Instructions with Human Feedbackâ. In: Advances in Neural Information Processing Systems. Vol. 35. Pechmann, Astrid, König, Kirsten, Bernert, GĂŒnther, Schachtrup, Kristina, Schara, Ulrike, Schorling, David, Schwersenz, Inge, Stein, Sabine, Tassoni, Adrian, Vogt, Sibylle, Walter, Maggie C., LochmĂŒller, Hanns, and Kirschner, Janbernd (2019). âSMArtCARE - a platform to collect real-life outcome data of patients 19 with spinal muscular atrophyâ. In: Orphanet J Rare Dis 14.1, p. 18. DOI: 10. 1186/s13023-019-0998-4. SchĂ€chter, Clemens, Hackenberg, Maren, Pfaffenlehner, Michelle, Tambe-Ndonfack, FĂ©lix B., Schmidt, Thorsten, Pechmann, Astrid, Kirschner, Janbernd, Hase- nauer, Jan, and Binder, Harald (2025). âUsing latent representations to link disjoint longitudinal data for mixed-effects regressionâ. In: arXiv preprint. DOI: 10.48550/arXiv.2510.25531. Schorling, David C., Pechmann, Astrid, and Kirschner, Janbernd (2020). âAdvances in Treatment of Spinal Muscular Atrophy â New Phenotypes, New Challenges, New Implications for Careâ. In: Journal of Neuromuscular Diseases 7.1, p. 1â 13. DOI: 10.3233/JND-190424. Singhal, Karan, Azizi, Shekoofeh, Tu, Tao, Mahdavi, S. Sara, Wei, Jason, Chung, Hyung Won, Scales, Nathan, Tanwani, Ajay, Cole-Lewis, Heather, Pfohl, Stephen, Payne, Perry, Seneviratne, Martin, Gamble, Paul, Kelly, Chris, Babiker, Abubakr, SchĂ€rli, Nathanael, Chowdhery, Aakanksha, Mansfield, Philip, Demner-Fushman, Dina, Arcas, Blaise AgĂŒera y, Webster, Dale, Corrado, Greg S., Matias, Yossi, Chou, Katherine, Gottweis, Juraj, TomaĆĄev, Nenad, Liu, Yun, Rajkomar, Alvin, Barral, Joelle, Semturs, Christopher, Karthikesalingam, Alan, and Natarajan, Vivek (2023). âLarge language models encode clinical knowledgeâ. In: Nature 620, p. 172â180. DOI: 10.1038/s41586-023-06291-2. Singhal, Karan, Tu, Tao, Gottweis, Juraj, et al. (2025). âToward expert-level med- ical question answering with large language modelsâ. In: Nature Medicine 31, p. 943â950. DOI: 10.1038/s41591-024-03423-7. Smith, Ryan, Fries, Jason A., Hancock, Braden, and Bach, Stephen H. (2022). âLanguage Models in the Loop: Incorporating Prompting into Weak Supervi- sionâ. In: arXiv preprint. DOI: 10.48550/arXiv.2205.02318. Sohn, Kihyuk, Lee, Honglak, and Yan, Xinchen (2015). âLearning structured out- put representation using deep conditional generative modelsâ. In: Advances in Neural Information Processing Systems 28, p. 3483â3491. URL: https:// papers.nips.c/paper/5775-learning-structured-output-representation- using-deep-conditional-generative-models. Tamuz, Omer, Liu, Ce, Belongie, Serge, Shamir, Ohad, and Kalai, Adam Tauman (2011). âAdaptively Learning the Crowd Kernelâ. In: arXiv preprint arXiv:1105.1033. arXiv: 1105.1033 [cs.LG]. Thirunavukarasu, Arun James, Ting, Darren Shu Jeng, Elangovan, Kabilan, Gutier- rez, Laura, Tan, Ting Fang, and Ting, Daniel Shu Wei (2023). âLarge language models in medicineâ. In: Nature Medicine 29, p. 1930â1940. DOI: 10.1038/ s41591-023-02448-8. 20 Thurstone, L. L. (1927). âA Law of Comparative Judgmentâ. In: Psychological Re- view 34.4, p. 273â286. DOI: 10.1037/h0070288. Varone, Antonio, Esposito, Gabriella, and Bitetti, Ilaria (2025). âSpinal Muscular At- rophy in the Era of Newborn Screening: How the Classification Could Changeâ. In: Frontiers in Neurology 16, p. 1542396. DOI: 10.3389/fneur.2025.1542396. Veen, Dave Van, Uden, Cara Van, Blankemeier, Louis, Delbrouck, Jean-Benoit, Bluethgen, Christian, Pareek, Anuj, Reis, Eduardo Pontes, Langlotz, Curtis P., Gatidis, Sergios, and Chaudhari, Akshay S. (2024). âAdapted large language models can outperform medical experts in clinical text summarizationâ. In: Na- ture Medicine 30, p. 1134â1142. DOI: 10.1038/s41591-024-02855-5. Vickers, Andrew J. and Elkin, Elena B. (2006). âDecision curve analysis: A novel method for evaluating prediction modelsâ. In: Medical Decision Making 26.6, p. 565â574. DOI: 10.1177/0272989X06295361. Yang, An, Li, Anfeng, Yang, Baosong, Zhang, Beichen, Hui, Binyuan, Zheng, Bo, Yu, Bowen, Gao, Chang, Huang, Chengen, Lv, Chenxu, et al. (2025). âQwen3 Technical Reportâ. In: arXiv preprint arXiv:2505.09388. 21 Encoder Decoder 2) Pairwise questions and final label assignment1) Prompt rendering Given a patient observed at time with profile and context , assign label or . Surrogate Network (frozen) â â ⚯ â â Latent space A â Of fline surrogate construction B â Synthetic-expert-supervised model fitting Clinical Data ContextObservation 3) Surrogate training Surrogate Network 4) Encode data and model latent dynamics 5) Reconstruct data and pass to surrogate 6) Align surrogate distributions via loss term 1. or 2. or 3. or 4. or 5. or 6. or 1. 2. 3. 4. 5. 6. LLM Clinical Labels Figure 1: Overview of the proposed synthetic-expert supervision framework. Part A: Offline surrogate construction. 1) Each observed clinical profile y i,t and its contextual information s i,t are rendered as a natural-language prompt. 2) The LLM rater produces pairwise responses for the specified candidate labels, which are checked for consistency and aggregated into retained synthetic-expert labels. 3) The retained LLM-derived labels are used to train a differentiable surrogate net- work S Ï that maps clinical profiles and contextual information to distributions over the clinical labels. Part B: Synthetic-expert-supervised model fitting. 4) The encoder maps the longi- tudinal observations to low-dimensional latent representations, whose trajectories are linked through the latent mixed-effects model. 5) The decoder reconstructs the clinical profiles, and both the observed profiles y i,t and their reconstructions Ë y i,t are evaluated by the frozen surrogate network using the same contextual in- formation s i,t . The third observation is reconstructed with the wrong label. 6) The JensenâShannon divergence between the corresponding surrogate distributions is added to the training objective as a synthetic-expert loss term, thereby discourag- ing reconstructions whose clinical interpretation differs from that of the observed profile. Alt text: Graphical approach visualization divided into two stages. Part A shows clinical observations and contextual information being converted into text prompts. An LLM performs pairwise comparisons between clinical labels, inconsistent re- sponses are identified, and the retained labels are used to train a differentiable surrogate network. Part B shows an encoder, longitudinal mixed-effects model, and decoder producing reconstructed clinical profiles. A frozen surrogate evalu- ates the original and reconstructed profiles, and the JensenâShannon divergence between their predicted label distributions is added to the training loss. 22 0123456 Age in years 0.0 0.2 0.4 0.6 0.8 1.0 Can sit, but not stand (Sit = 4, Stand < 2) 0123456 Age in years Can stand, but not walk (Stand = 3, Walk < 2) 0123456 Age in years Can walk (Walk = 3) Predicted class probability Synthetic expert class probabilities patients with sit, stand and walk milestones SMA type SMA1 SMA2 SMA3 Presymptomatic SMA type SMA1 SMA2 SMA3 Presymptomatic Synthetic expert Qwen-3-14B GPT-OSS-120B Gemma-4-31B (a) Age-dependent surrogate class probabilities for selected motor-function profiles. â1.0â0.50.00.51.0 z1 â1.0 â0.5 0.0 0.5 1.0 z2 Gemma-4-31B â1.0â0.50.00.51.0 z1 GPT-OSS-120B â1.0â0.50.00.51.0 z1 Qwen3-14B SMA1 SMA2 SMA3 Presymptomatic Latent space colored by expected SMA type at age = 1 years (b) Latent-space regions colored by surrogate-predicted SMA type. Figure 2: Interpretability of the synthetic-expert surrogate in profile space and la- tent space. The upper panel shows age-dependent SMA type probabilities for three patient groups: sitting without standing (observations with HINE-2 sitting item score = 4, standing < 2), standing without walking (standing = 3, walking < 2), and walking (walking = 3, corresponds typically with a full HINE-2 sum score with no recorded deficiencies). Colors indicate SMA type labels and marker shapes indicate the LLM model. The lower panel visualizes the latent space at age = 1 year colored by different synthetic experts. Each point in the latent grid is decoded into a motor-function profile and then classified by the selected LLM-specific surrogate network. Alt text: Composite figure with two rows of three panels. Panel (a) shows pre- dicted probabilities for four SMA-related categories as functions of age for profiles characterized by sitting without standing, standing without walking, and indepen- dent walking. Curves compare the Gemma-4-31B, GPT-OSS-120B, and Qwen-3- 14B surrogate models. Sitting-only profiles generally shift from presymptomatic at very young ages toward SMA type 2; standing profiles shift toward SMA types 2 or 3; and walking profiles remain predominantly presymptomatic. Panel (b) shows category regions across two latent dimensions at age one year for the same three models. All models produce broadly similar severity-ordered regions, although their category boundaries differ locally. 23 Table 1: Diagnostics of LLM-derived synthetic-expert judgments and surrogate training. Ordinal consistency is the percentage of profiles with internally consistent pairwise answers. Pairwise and final-label agreement are measured under label- order reversal. Final-label agreement is defined as when the final label agrees un- der both prompt presentation orders, or the pairwise questions completely align for inconsistent answer patterns. Surrogate accuracy is the cross-validated accuracy for reproducing retained synthetic-expert labels. Values for the coin toss baseline are analytical values. The uncertainty ratio is the mean self-reported uncertainty for pairwise answers that changed under label-order reversal, divided by the mean for answers that remained stable. Here, low was mapped to an uncertainty score of 0, medium to 1 and high to 2. Synthetic expert Ordinal consistency (%) Rev. pairwise agreement (%) Rev. final-label agreement (%) Surrogate accuracy (%) Uncertainty ratio Coin Toss9.450.01.733.3â Gemma-4-E2B74.185.352.889.61.5 Gemma-4-E4B80.084.647.884.81.3 Gemma-4-12B87.195.785.194.911.4 Gemma-4-31B89.997.892.797.29.6 GPT-OSS-20B74.185.848.787.51.6 GPT-OSS-120B82.090.263.591.82.2 Qwen-3-0.6B10.958.16.263.41.0 Qwen-3-1.7B34.768.614.968.71.0 Qwen-3-4B48.683.842.691.11.4 Qwen-3-8B80.791.870.096.12.1 Qwen-3-14B82.193.676.592.42.2 Table 2: Effect of synthetic-expert supervision weight on reconstruction consis- tency and reconstruction error. Lower values indicate better consistency or lower reconstruction error. Label disagreement is averaged over the forward and re- verse label presentation orders. Bold values indicate the best value within each synthetic-expert model class and metric. Synthetic expert λ SE JS divergence Label disagreement (%) HINE-2 sum-score error HINE-2 error among mismatches Gemma-4-31B00.05410.91.3592.451 100.0439.01.3572.012 250.0388.21.3602.010 500.0347.51.3772.004 1000.0316.81.3822.001 GPT-OSS-120B00.03410.11.3592.354 100.0298.71.3671.924 250.0268.11.3581.910 500.0247.71.3611.889 1000.0227.21.3701.872 Qwen-3-14B00.0328.91.3592.351 100.0277.91.3562.005 250.0257.31.3451.898 500.0226.71.3541.827 1000.0216.41.4251.853 24 Table 3: Milestone prediction on latent representations influenced by synthetic- expert supervision. Rows report one-year landmark prediction performance for a linear time-varying Cox model. Brier denotes the inverse-probability-of-censoring weighted Brier score, AUC the corresponding time-dependent AUC, and ECE the weighted expected calibration error across calibration bins. Lower Brier and ECE values and higher AUC values indicate better performance. Bold values indi- cate the best performance across all latent-representation feature sets for each milestone-metric combination. SittingStandingWalking Feature setλ SE BrierAUCECEBrierAUCECEBrierAUCECE HINE-2 sum scoreâ0.271 0.817 0.314 0.180 0.834 0.175 0.142 0.845 0.134 HINE-2 sum score + covariates â0.209 0.890 0.255 0.155 0.908 0.153 0.124 0.919 0.119 Z00.097 0.944 0.078 0.098 0.951 0.102 0.087 0.954 0.088 Z informed by Gemma-4-31B 100.085 0.954 0.067 0.084 0.962 0.076 0.067 0.964 0.065 250.084 0.956 0.064 0.088 0.960 0.081 0.071 0.962 0.067 500.085 0.952 0.058 0.089 0.956 0.079 0.070 0.962 0.066 1000.091 0.951 0.075 0.094 0.954 0.084 0.074 0.960 0.068 Z informed by GPT-OSS-120B 100.086 0.952 0.061 0.079 0.962 0.063 0.064 0.967 0.065 250.083 0.954 0.057 0.082 0.964 0.073 0.066 0.967 0.063 500.091 0.958 0.090 0.096 0.959 0.092 0.077 0.959 0.065 1000.098 0.955 0.108 0.096 0.956 0.087 0.076 0.957 0.072 Z informed by Qwen-3-14B 100.094 0.951 0.070 0.089 0.962 0.096 0.078 0.967 0.080 250.091 0.952 0.072 0.089 0.958 0.082 0.075 0.958 0.071 500.096 0.949 0.071 0.091 0.958 0.091 0.083 0.961 0.084 1000.102 0.939 0.077 0.097 0.955 0.096 0.081 0.959 0.087 25 Supplementary materials for: Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling Clemens SchĂ€chter 1,2 , Astrid Pechmann 3 , Janbernd Kirschner 3 , and Harald Binder 1,2,4 1 Institute of Medical Biometry and Statistics (IMBI), Faculty of Medicine and Medical Center â University of Freiburg, Freiburg, Germany 2 Freiburg Center for Data Analysis, Modeling and AI â University of Freiburg, Freiburg, Germany 3 Department of Neuropediatrics and Muscle Disorders, Faculty of Medicine and Medical Center â University of Freiburg, Freiburg, Germany 4 CIBSS, Centre for Integrative Biological Signalling Studies â University of Freiburg, Freiburg, Germany A Data preprocessing and hyperparameters A.1 Dataset The analysis used HINE-2 motor-function assessments of patients under the age of 12 years from the SMArtCARE registry. The SMArtCARE registry tracks the motor function development, motor milestone achievement and treatment history of children diagnosed with spinal muscular atrophy. The HINE-2 profile con- tains eleven motor-function items regarding upper limb function, raising hands, reaching overhead, head control, sitting, voluntary grasp, kicking, rolling, crawl- ing, standing, and walking. The measured motor abilities of these categorical item responses were mapped to ordered numerical levels, and visits with missing values were excluded. For neural-network inputs, each ordinal HINE-2 item was subsequently thermometer-encoded. To retain longitudinal information for latent mixed-effects modeling, patients were required to have more than four complete HINE-2 visits. 1 In total the dataset contained exactly 13,000 observations and 994 patients, with a median of 12 visits per patient. The median age at first observation was 1.2 years, and the median follow-up duration was 3.5 years. A.2 Synthetic-expert label generation and surrogate training To generate a training dataset for the surrogate network we supplemented the real HINE-2 data with synthetic samples. For this we fitted a two-dimensional age- conditional VAE model to the HINE-2 profiles. Subsequently, synthetic profiles were sampled from the decoder until a total of 20,000 unique generated profiles were retained. To allow for slight variations in ages, the age values were jittered by up to 0.25 years. Each profile was rendered as a textual motor-function description together with the age value. The item-level wording was based on the HINE-2 item de- scriptions [1]. The exact score-to-text conversion used for this rendering is shown in Table S1. Table S1: Text rendering of HINE-2 profile scores used for the synthetic-expert prompts. Each generated profile was converted by selecting the sentence corre- sponding to the observed item score and concatenating the resulting item-level descriptions with the patientâs age. HINE-2 item Text fragment by item score IntroductionThe child is years years months months old when its motor function is evaluated. The child has the following motor function skills: Upperlimb function 0: Cannot achieve useful hand function to grab objects. 1: Achieves useful hand function to grab objects. Raise hands 0: Cannot raise its hands to the mouth. 1: Raises its hands to the mouth. Reach over- head 0: Cannot reach overhead. 1: Reaches overhead. Head control 0: Has no stable head control and cannot maintain the head upright. 1: Holds the head upright briefly, but the head wobbles and control is unsteady. 2: Maintains the head upright consistently and does not show head wobbling. Sitting0: Cannot maintain a sitting position in a meaningful way even with external support. 1: Sits only if the examiner stabilizes the pelvis or hips and cannot sit independently. 2: Sits with propping and uses the arms or hands for balance, but cannot sit freely without upper limb support. 3: Sits independently without obvious propping, but cannot rotate or pivot while sitting. 4: Sits independently and rotates and pivots while sitting. Voluntary grasp 0: Does not show a voluntary grasp and cannot intentionally take or hold an object. 1: Grasps using the whole hand, but cannot perform a more refined finger-thumb grasp or pincer grasp. 2: Grasps using index finger and thumb in an immature grasp, but cannot perform a pincer grasp. 3: Performs a pincer grasp and picks up objects with a mature finger-thumb grip. Continued on next page 2 Table S1: Text rendering of HINE-2 profile scores used for the synthetic-expert prompts (continued). HINE-2 item Text fragment by item score Kicking0: Cannot kick in the supine position. 1: Kicks horizontally in the supine position, but cannot lift the legs upwards. 2: Kicks upward vertically in the supine position, but cannot reach the legs with the hands. 3: Lifts the legs and touches the leg in the supine position, but cannot reach the toes. 4: Lifts the legs in the supine position high enough to touch the toes. Rolling0: Cannot roll onto the side from supine. 1: Rolls onto the side from supine, but cannot complete rolling between supine and prone. 2: Rolls from prone to side and supine, but cannot roll from supine to prone. 3: Rolls fully in both directions, including from supine to prone and vice versa. Crawling0: Cannot lift the head in a way that supports crawling. 1: Supports weight on the elbows, but cannot push up on outstretched hands or crawl. 2: Supports weight on outstretched hands, but cannot crawl forward. 3: Crawls flat on the abdomen, but cannot crawl on hands and knees. 4: Crawls on hands and knees. Standing0: Cannot bear weight through the legs when held upright. 1: Bears weight through the legs when held upright, but cannot stand with external support. 2: Stands with external support, but not unaided. 3: Stands upright, stable and unaided. Walking0: Cannot show walking-related motor behavior and cannot bounce, cruise, or walk independently. 1: Shows early supported stepping behavior, but cannot cruise along furniture or walk independently. 2: Cruises while holding on for support, but cannot walk independently. 3: Walks independently for at least several seconds. Each profile was then queried through pairwise comparisons between the can- didate labels SMA type 1, SMA type 2, SMA type 3, and presymptomatic motor development. Additionally, the pairwise comparisons were repeated under label- order reversal, resulting in two comparison panels per profile. The prompt tem- plate used for each pairwise comparison is shown in Table S2. Table S2: Prompt template for one pairwise synthetic-expert comparison. The placeholders label A, label B, and profile text were filled for each candidate-label pair and generated HINE-2 profile. Prompt partText Comparison instructionDecide between two candidate motor-function labels for a child being evaluated for spinal muscular atrophy (SMA): label A or label B. Decision ruleUse the childâs age to contextualize the observed motor abilities. If neither label is a rea- sonable match, choose the allowed label whose severity is closer to the observed motor impairment. **When label A or label B is presymptomatic the following additional in- struction was included:** Choose presymptomatic when motor abilities are age-appropriate and no SMA-related motor impairment is observed. Profile textprofile text Output constraintThink briefly about your answer. Report your confidence of your final decision and con- clude by returning the following JSON object: "label": "label B", "certainty": "high|medium|low" or "label":"label A", "certainty": "high|medium|low". The pairwise answers were aggregated into final SMA type labels by using 3 the ordinal structure of the four labels, SMA type 1 < SMA type 2 < SMA type 3 < presymptomatic motor development. Response patterns containing cycles or non-monotone preferences were classified as internally contradictory. Profiles for which the label order failed to produce a valid final label were excluded from surrogate training. The differentiable surrogate was then trained as a classifier to reproduce the retained labels from patient age and the thermometer-encoded HINE-2 profile. Training was performed with two heads to reproduce the answer for each label order. Results were averaged over both heads. A.3 Hyperparameters All neural network architectures were chosen to contain a single hidden layer with 100 units and ReLU activation function. Thermometer encoded HINE-2 profiles together with the age value amounted to 41 input units to the encoder network and surrogate model. The surrogate was trained with cross-entropy loss on class labels, using the Adam optimizer with default hyperparameters and 50 training epochs. The conditional VAE used a two-dimensional latent space and received age as conditioning information. The decoder parameterized an ordinal categorical likeli- hood for the HINE-2 items. cVAE parameters were optimized with Adam a learn- ing rate of 0.01 over 20 alternating training epochs. The KL term was weighted by ÎČ = 0.5, the alignment term by η = 5. The synthetic-expert supervision weight was varied over λ SE â 0, 10, 25, 50, 100. The latent mixed-effects model was fitted in the same two-dimensional latent space. For determining fixed-effect parameters, we followed our previous work [2] and included fixed effects for age at symptom onset, SMN2 copy-number, time since treatment for the three recorded SMA medications, time since switching to a different SMA medication, BMI, height, and interaction terms with age. Treatment- time variables were set to zero before the respective treatment or switch. The random-effects structure contained a patient-specific random intercept and ran- dom slope. The mixed-effects parameters were re-estimated after each cVAE epoch with L-BFGS, using a learning rate of 0.1. A.4 Milestones and Cox regression The SMArtCARE registry records patient-level attainment of independent sitting, standing, and walking as motor milestones. These endpoints correspond to the same broad motor domains as the HINE-2 items, but use separate milestone 4 definitions. For the Cox analysis, exact milestone-attainment ages were treated as uncensored events. If follow-up information was available but the exact at- tainment time was not observed, patients were right-censored at the recorded age. If future milestone attainment was unknown, they were right-censored at the last motor-function visit without a recorded milestone. Entries indicating that a milestone had already been achieved before observation, but without an exact age, were left-censored and excluded. The milestone definitions and number of observed events are shown in Table S3. Table S3: Definitions of the motor milestones used for Cox regression. Events denote exact milestone attainments contributing an event interval after excluding left-censored patient-milestone pairs and pairs without a motor-function visit be- fore the event time. Milestone Events Definition Sitting432 Patient sits up straight with the head erect for at least 10 seconds. Patient does not use arms or hands to balance the body or to support the position. Standing304 Patient stands in an upright position on both feet, not on the toes, with the back straight. The legs support 100% of the patientâs weight. There is no contact with an object or a person. Patient stands alone for at least 10 seconds. Walking264 Patient takes at least 5 steps independently in an upright position with the back straight. One leg moves forward while the other supports most of the body weight. There is no contact with a person or object. For each patient and milestone, the longitudinal visits before the event or cen- soring time were converted into time-varying Cox intervals on the age scale. Each interval started at an observed visit age and ended at the next visit age or at the milestone or censoring time, whichever occurred first. The evaluated feature sets comprised the HINE-2 total score, the HINE-2 total score plus fixed clinical co- variates and latent encoder coordinates. For the synthetic-expert supervision sweep, linear Cox models were fitted to the latent representations across syn- thetic experts with different choices of supervision weights. Models were fitted with patient-grouped five-fold cross-validation. Predicted risks were evaluated at a fixed one-year landmark horizon from each held-out visit. For this horizon, a visit was treated as a case if the mile- stone occurred within one year and as a control if the patient remained event-free beyond the horizon. Visits censored before the one-year horizon were handled through inverse-probability-of-censoring weighting. Evaluation rows were addi- tionally weighted inversely by the number of landmark visits per patient and mile- stone, so that patients with more visits did not dominate the metric calculation. We reported three one-year prediction metrics. Discrimination was measured by the inverse-probability-of-censoring weighted time-dependent AUC. Prediction error was measured by the corresponding weighted Brier score. Calibration was 5 summarized by the expected calibration error, defined as the weighted mean ab- solute difference between predicted and observed risk across 10 quantile-based calibration bins. 6