Paper deep dive
RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation
Ganlin Feng, Yuxi Long, Hafsa Ali, Erin Lou, Fahad Butt, Qian Liu, Yang Wang, Pingzhao Hu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 4/10/2026, 2:10:10 AM
Summary
RDFace is a curated benchmark dataset for pediatric rare disease facial analysis, containing 456 images across 103 conditions. It addresses extreme data scarcity by providing a standardized evaluation protocol for supervised and few-shot learning, and explores synthetic data augmentation using DreamBooth and FastGAN to improve diagnostic accuracy.
Entities (5)
Relation Signals (3)
RDFace â contains â Rare Disease
confidence 95% · RDFace, a curated benchmark dataset comprising 456 pediatric facial images spanning 103 rare genetic conditions
DreamBooth â augments â RDFace
confidence 90% · explore synthetic augmentation with DreamBooth and FastGAN
Qwen2.5-VL â evaluates â RDFace
confidence 85% · we use the Qwen2.5-VL... to generate diagnostic-style clinical reports from both real patient photos and DreamBooth-generated synthetic images
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Rare diseases often manifest with distinctive facial phenotypes in children, offering valuable diagnostic cues for clinicians and AI-assisted screening systems. However, progress in this field is severely limited by the scarcity of curated, ethically sourced facial data and the high similarity among phenotypes across different conditions. To address these challenges, we introduce RDFace, a curated benchmark dataset comprising 456 pediatric facial images spanning 103 rare genetic conditions (average 4.4 samples per condition). Each ethically verified image is paired with standardized metadata. RDFace enables the development and evaluation of data-efficient AI models for rare disease diagnosis under real-world low-data constraints. We benchmark multiple pretrained vision backbones using cross-validation and explore synthetic augmentation with DreamBooth and FastGAN. Generated images are filtered via facial landmark similarity to maintain phenotype fidelity and merged with real data, improving diagnostic accuracy by up to 13.7% in ultra-low-data regimes. To assess semantic validity, phenotype descriptions generated by a vision-language model from real and synthetic images achieve a report similarity score of 0.84. RDFace establishes a transparent, benchmark-ready dataset for equitable rare disease AI research and presents a scalable framework for evaluating both diagnostic performance and the integrity of synthetic medical imagery.
Tags
Links
- Source: https://arxiv.org/abs/2604.03454v1
- Canonical: https://arxiv.org/abs/2604.03454v1
Trouble viewing inline? Open PDF directly â
Full Text
94,582 characters extracted from source content.
Expand or collapse full text
RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation Ganlin Feng 1 Yuxi Long 1 Hafsa Ali 2 Erin Lou 3 Fahad Butt 1 Qian Liu 4 Yang Wang 2 Pingzhao Hu 1,3 * 1 Western University 2 Concordia University 3 University of Toronto 4 University of Winnipeg Abstract Rare diseases often manifest with distinctive facial phe- notypes in children, offering valuable diagnostic cues for clinicians and AI-assisted screening systems.However, progress in this field is severely limited by the scarcity of curated, ethically sourced facial data and the high similar- ity among phenotypes across different conditions. To ad- dress these challenges, we introduce RDFace, a curated benchmark dataset comprising 456 pediatric facial images spanning 103 rare genetic conditions (average 4.4 sam- ples per condition). Each ethically verified image is paired with standardized metadata. RDFace enables the develop- ment and evaluation of data-efficient AI models for rare disease diagnosis under real-world low-data constraints. We benchmark multiple pretrained vision backbones using cross-validation and explore synthetic augmentation with DreamBooth and FastGAN. Generated images are filtered via facial landmark similarity to maintain phenotype fi- delity and merged with real data, improving diagnostic ac- curacy by up to 13.7% in ultra-low-data regimes. To as- sess semantic validity, phenotype descriptions generated by a visionâlanguage model from real and synthetic im- ages achieve a report similarity score of 0.84. RDFace es- tablishes a transparent, benchmark-ready dataset for eq- uitable rare disease AI research and presents a scalable framework for evaluating both diagnostic performance and the integrity of synthetic medical imagery. Project page: https://github.com/Kkathyf/RDFace. 1. Introduction Rare diseases (RDs) affect only a limited fraction of indi- viduals in the general population compared to more preva- lent diseases. To date, more than 10,000 RDs have been identified, each with a prevalence of 1 in 2000 or less [15]. Although these disorders are rare individually, nearly 350 million people are affected by RDs worldwide [10]. The * Corresponding author (phu49@uwo.ca) heterogeneity in symptoms, wide range of disorders, limited data, and geographic dispersion all make RD a challenging domain for research [6]. Traditional diagnostic pathways for pediatric RDs are of- ten complex and time-consuming, typically involving mul- tiple clinical assessments, genetic testing, and specialist re- ferrals [55]. As a result, patients frequently experience pro- longed diagnostic odysseys before reaching a definitive di- agnosis [4], delaying appropriate treatment and imposing emotional and financial burdens on families and healthcare systems [1, 55]. These barriers are further amplified in ru- ral or underserved regions, where access to specialized care and timely diagnosis is especially limited [42]. A RD Eu- rope study across 17 European countries revealed that 25% of patients with RDs, such as Duchenne muscular dystro- phy and Ehlers-Danlos syndrome, were not properly diag- nosed until 5 to 30 years after symptom onset [9]. These challenges are especially critical in childhood since many RDs share overlapping symptoms with common childhood conditions, making early clinical recognition difficult [4]. Notably, many genetic syndromes manifest distinctive craniofacial phenotypes in early childhood, making facial analysis a promising non-invasive signal for RD recogni- tion [27, 38]. Recent advances in computer vision have demonstrated the potential of facial image analysis for as- sisting clinical screening and syndrome identification [14]. However, progress still remains limited by the lack of stan- dardized benchmarks that reflect realistic clinical scenarios, particularly under extreme data scarcity where only a few samples per condition are available. To address this gap, we introduce RDFace, a curated benchmark dataset for pediatric RD facial recognition under extreme low-shot conditions. RDFace provides a unified evaluation protocol for systematically comparing learning strategies and analyzing phenotype-aware synthetic aug- mentation. We consider on three key aspects: (i) bench- mark construction under realistic low-data settings, (i) the effectiveness of phenotype-aligned synthetic augmentation, and (i) the assessment of generated image fidelity via arXiv:2604.03454v1 [cs.CV] 3 Apr 2026 Figure 1. Pipeline of RDFace Evaluation. The pipeline includes dataset curation, classification tasks, synthetic data generation and augmentation, and benchmark evaluation for RD diagnostic support. phenotype-level similarity. The overall evaluation pipeline is illustrated in Figure 1. Overall, our contributions are summarized as follows: âą Benchmark Dataset: We introduce RDFace, a curated pediatric RD facial image benchmark containing 456 im- ages across 103 conditions, where each class contains only 1â7 samples. âą Synthetic Augmentation Study: Using a unified bench- mark protocol, we analyze the effect of phenotype- aligned synthetic facial image augmentation in ultra-low- shot recognition. âą Phenotype-based Fidelity Evaluation: We propose a phenotype-based evaluation using vision-language mod- els (VLMs) to compare phenotype descriptions from real and synthetic images. 2. Related work Facial morphology has long supported diagnosis in syn- dromic disorders [26, 40].Advances in deep learning have enabled automated facial phenotyping, bridging ge- netic medicine and computer vision [11, 12, 49]. Deep learning models have been used to classify genetic syn- dromes from facial images, but typically focus on under 15 well-represented syndromes with hundreds of images per class [23, 44]. While GestaltMatcher extends to hundreds of syndromes via retrieval-based matching [20], it still re- quires enough training samples per condition, which is of- ten unrealistic for ultra-RDs. To cope with data scarcity, few-shot learning has gained traction as an effective strategy [31]. Metric-based methods like Prototypical Networks [46] show promise in medical imaging [34], but remain underexplored for RD facial diag- nosis due to high intra-class variability and subtle inter-class differences. Synthetic data generation offers another strat- egy for augmenting low-resource datasets. While Genera- tive Adversarial Network (GAN) [13] and diffusion mod- els [19] have shown success in other medical contexts [22], they often require large training sets and fine-tuning. Fast- GAN improves low-data synthesis via skip-layer excitation and self-supervision [29], while DreamBooth is capable of generating high-fidelity samples from 3â5 images [41]. In the RD setting, GestaltGAN [25] enables face synthesis for privacy, but does not evaluate its synthetic images in down- stream diagnostic tasks. Recently, multimodal VLMs have shown promise in linking medical images with clinical semantics for report generation and phenotype interpretation [17]. Such models have been applied in radiology and pathology [50] to ex- plain predictions and assess imageâreport consistency, but remain underexplored in RD phenotyping. 3. Dataset description 3.1. Data collection and sources The data collection process was conducted in accordance with institutional research ethics guidelines and received approval from the appropriate Research Ethics Board, en- suring ethical use of publicly sourced clinical imagery. We curated a dataset of 456 frontal pediatric facial portraits spanning 103 rare genetic diseases. Detailed information about each disease is provided in Appendix A.1. Candi- date conditions were identified from clinical observations in populations with elevated rates of monogenic disorders, often due to historical founder effects and genetic isolation. Although regionally concentrated, many of these diseases are globally recognized and characterized by distinctive craniofacial features, enhancing their relevance for broader a b Average Age: 6.36 c a Figure 2. Overview of RDFace Dataset. (a) Geographic distribution of cases across 46 countries, with circle size proportional to the number of cases per country. (b) Sex distribution. (c) Age group distribution of patients with an average age of 6.36 years. research. Each condition was cross-verified with Orphanet [48] to confirm its rarity. For each disease, we recorded metadata including gene association (if known) and a stan- dardized abbreviation for labeling and analysis. To construct the dataset, we targeted 1-5 high-quality images per disease class, focusing on patients aged 0â18, with emphasis on children under 12 to minimize visual confounding from adult comorbidities [3]. Eligible images were required to be frontal-facing, portrait-style, and of suf- ficient resolution, showing a single patient with open eyes. Images were sourced through structured web searches us- ing disease names and terms such as âface,â âchild,â and âpatient.â We prioritized peer-reviewed literature, hospi- tal foundations, and verified clinical reports, and included manually reviewed advocacy content when necessary. An overview of the demographic diversity of RDFace is shown in Figure 2. More dataset details are in Appendix A.2. Expert review. The dataset curation, creation, and overall design were conducted under the supervision of clinical ge- neticists to ensure medical plausibility and alignment with diagnostic practice. The real images were initially verified by their original sources, with validation by clinicians ex- plicitly stated in the associated publications. To further en- sure dataset reliability, two clinical fellows independently reviewed the plausibility of the image-label associations. 3.2. Dataset structure RDFace is structured as a class-organized image dataset de- signed for pediatric rare disease classification. The dataset is stored in a directory named rd images/, which contains 103 subfolders, each corresponding to a rare disease class. Each subfolder is named using the disease abbreviation and contains images representing pediatric patients diagnosed with file names [disease abbr].[index].png. To facilitate dataset navigation and programmatic access, a metadata file named diseaseimages.csv is provided in the root directory. This file contains image-level annotations across the follow- ing fields: (1) image name, (2) disease name, (3) gene, (4) disease abbreviation, (5) disease subcategory, and (6) Or- phanet code. Demographic attributes are summarized at the dataset level but not shared as metadata due to inconsistent availability. The disease abbreviation serves as the class la- bel for all classification tasks conducted in this work. 4. Methodology This section outlines the benchmarking methodology used in RDFace, encompassing four complementary compo- nents: baseline supervised classification, few-shot evalua- tion, phenotype-aware synthetic data generation, and down- stream analysis of synthetic data utility. 4.1. Standard supervised classification The RDFace dataset is split into training and testing par- titions using a stratified 75%/25% image-level split to ap- proximately preserve class distributions. Singleton classes (i.e., those with only one sample) are assigned entirely to the training set to ensure sufficient representation and avoid missing classes during evaluation. During training pro- cess, 5-fold cross-validation is performed for hyperparam- eter tuning, with folds stratified and singleton samples al- ways retained in the training folds. We benchmark six pretrained backbone architectures: ResNet-152 [18], DenseNet-169 [21], FaceNet [43], VGG- 16 [45], Swin Transformer [32], and CLIP (ViT-B/32) [39], using weights initialized from ImageNet [7], VG- GFace [35], or VGGFace2 [5]. Input images are resized to 224Ă 224 and normalized with ImageNet statistics. The final classification layer is replaced with a 103-way soft- max corresponding to disease classes. Models are trained using only real training images and evaluated on the held- out test set using Top-k accuracy metrics, where a predic- tion is considered correct if the true label appears among the modelâs top k ranked outputs. Top-k accuracy is the stan- dard evaluation protocol in facial phenotype recognition and rare disease classification benchmarks, ensuring compara- bility with prior studies. External architecture comparison. For external refer- Figure 3. Pipeline of synthetic data generation and evaluation. Real pediatric facial images are first preprocessed using Real-ESRGAN and DDColor, then used to generate synthetic faces via DreamBooth (class-conditioned) and FastGAN (unconditional). Generated images are evaluated for facial realism (RetinaFace and LPIPS) and phenotype consistency (landmark-based cosine similarity). ence, we take DeepGestalt as the representative architecture of prior diagnostic systems. We reproduced its backbone configuration (refered to as the Gestalt throughout this pa- per) using a classifier trained under the RDFace protocol. 4.2. Few-shot learning We evaluate few-shot learning on RDFace using Prototyp- ical Networks under an n-way 1-shot classification setting. Prototypical Network details are provided in Appendix B. Singleton classes are excluded, resulting in 99 RD classes. These classes are randomly split into 80% for training and 20% for testing, with no class overlap. Within the training set, 5-fold cross-validation is used by holding out a disjoint subset of classes for validation in each fold. Episodic training is used to match the few-shot evalua- tion protocol. In each episode, n distinct classes are sam- pled (e.g., n = 5), with one support and one query im- age per class. Prototypes are derived from support em- beddings, and queries are classified by their Euclidean dis- tance to these prototypes. The model is optimized using cross-entropy loss computed over query predictions, with gradients back-propagated to update the feature extractor. We train each fold for 600 episodes, using 100 additional episodes for validation. The same six pretrained backbones used in standard su- pervised classification are used to encode images. Final performance is evaluated on the meta-test set across three settings: 5-way 1-shot, 10-way 1-shot, and 15-way 1-shot. For each setting, we sample 200 test episodes from unseen classes and report mean accuracy averaged across the five cross-validation folds with standard deviation. 4.3. Synthetic data generation To further address the limited availability of real train- ing data in RD facial classification, we explore syn- thetic data augmentation through generative modeling. Two approaches are employed: DreamBooth, a super- vised diffusion-based model fine-tuned for each disease, and FastGAN, an unsupervised generative adversarial net- work. These models were selected for their complementary strengths: DreamBooth enables phenotype-aware genera- tion that captures class-specific facial features by condition- ing on text and exemplars, while FastGAN offers efficient training on limited data and generates class-agnostic sam- ples to promote overall data diversity. An overview of the generation and evaluation pipeline is illustrated in Figure 3. To improve image quality, real images are enriched through super-resolution to 512 Ă 512 pixels using Real- ESRGAN [47] and colorization using DDColor [24]. Five key facial landmarks (two eye centers, one nose tip, and two mouth corners) were extracted from preprocessed images using RetinaFace [8]. This five-point configuration is se- lected for its robustness and consistent detectability across varied image conditions and populations. DreamBooth. Each RD class is fine-tuned separately using its available preprocessed images in the training set, guided by the text prompt âa child with [disease abbr] diseaseâ. After fine-tuning, 100 synthetic images are generated per disease class. To ensure data quality, outputs flagged as âNSFWâ by the modelâs safety filter are removed. FastGAN. FastGAN is trained unconditionally from scratch on processed training images for 80,000 iterations, saving checkpoints every 10,000 iterations with 1,000 im- ages generated per checkpoint. 4.4. Synthetic data evaluation protocol Following prior work that jointly evaluates synthetic med- ical data for both perceptual realism and task-level fidelity [51], we adopt a dual evaluation framework encompassing (i) intrinsic realism and phenotype fidelity, and (i) down- stream diagnostic utility (Section 4.5). This section focuses on the first component, which evaluates the visual realism and disease relevance of synthetic images. Facial realism of synthetic images is evaluated via Reti- naFace detection scores and LPIPS perceptual similarity [54] related to real images. RetinaFace assigns a confidence score per image; scores above 0.90 are considered valid, and those exceeding 0.99 are deemed high-quality. Images are symmetrically padded during preprocessing to improve de- tection for tightly cropped faces. To assess the disease relevance of synthetic images, we employ a three-part evaluation protocol combining landmark-based similarity, expert clinical review, VLM- assisted phenotype interpretation. 4.4.1. Landmark-based similarity analysis We first adapt a landmark-based similarity framework to as- sess phenotypic similarity. For both real and synthetic im- ages, we compute a 5Ă5 pairwise Euclidean distance matrix from facial landmarks. An average distance matrix is then derived by aggregating real image matrices for each class. Each synthetic imageâs matrix is compared to these class- level averages using cosine similarity. Synthetic images are ranked by similarity to real class prototypes, with lower average rank indicating stronger alignment to true disease phenotypes. After validating this ranking on DreamBooth outputs, we pseudo-label FastGAN-generated images by as- signing each to the class with the highest similarity. 4.4.2. Expert review of synthetic samples We conducted an expert review to assess the clinical plau- sibility of synthetic images to complement automated eval- uations. A random subset of images generated by Dream- Booth and FastGAN was selected across a range of disease classes. Two medical doctors (MD) were invited to inde- pendently evaluate each image using a standardized eval- uation form. The primary criterion was visual consistency, whether the facial features in the synthetic image were plau- sible for the disease label provided. Experts were instructed to consider known clinical characteristics and phenotypic patterns, guided by paired real training images and their own medical knowledge. For each sample, they classified the image as either Plausible, Implausible, or Uncertain. To quantify inter-rater reliability, we computed evalua- tion metrics including the percentage of observed agree- ment, Cohenâs Îș coefficient, its standard error (SE), and the 95% confidence interval (CI). These metrics provide an objective measure of annotation consistency and serve as a quality control layer beyond automated scoring and embedding-based similarity analyses. 4.4.3. VLM-based phenotype description To supplement structural similarity analysis, we further in- troduce a structured evaluation protocol for assessing the clinical plausibility of synthetic facial images by leveraging the capabilities of VLMs. Specifically, we use the Qwen2.5- VL [2] and LLaVA-NeXT [30] models with a standardized prompt (Appendix G.1) to generate diagnostic-style clini- cal reports from both real patient photos and DreamBooth- generated synthetic images. The facial regions are deliberately chosen to align with the five facial landmarks used in our landmark-based simi- larity evaluation framework. The model is instructed to de- scribe observations in each region independently and con- clude with a predicted diagnosis and a clinical recommen- dation. This allows us to treat the VLMs as a form of se- mantic âphenotype readerâ, capable of interpreting real and synthetic facial morphology in clinical language. To quantify the similarity between phenotype descrip- tions, we compute semantic similarity using BioBERT [28] embeddings followed by cosine similarity, yielding region- wise and overall alignment scores for each pair. We evaluate similarity across three pair types: realâreal, realâsynthetic, and syntheticâsynthetic. Realâreal similar- ity serves as a reference for the consistency of phenotype descriptions from real data, realâsynthetic reflects the fi- delity of generated images to real disease characteristics, and syntheticâsynthetic provides insight into the consis- tency of generated samples. Uncertainty and robustness. To assess intra-model sta- bility, we perform a sampling-based uncertainty evaluation using real facial images from RDFace. For each image, the model generates 5 independent reports for each tempera- ture T â0.7, 0.9, 1.1 using stochastic decoding and then compute pairwise uncertainty 1 â mean similarity, where higher values indicate less consistent predictions. We fur- ther evaluate cross-model robustness by comparing pheno- type descriptions generated by Qwen2.5-VL and LLaVA- NeXT under the same protocol. 4.5. Downstream analysis of synthetic data utility As the second component of the dual evaluation framework, we assess the diagnostic utility of phenotype-aligned syn- thetic data by injecting them into standard supervised clas- sification and few-shot learning pipelines to quantify per- formance gains and scaling behavior. After generating synthetic samples, we evaluate their ef- fectiveness across standard supervised and few-shot learn- ing tasks. Within each disease class, synthetic images are ranked by landmark-based similarity to the real class proto- type. We then form augmentation sets of increasing scale (Top-n across all classes; n â 1000, 2000, 4000, 6000) to study scaling behavior. To verify that landmark-selected samples also maintain visual realism, we compute the aver- age RetinaFace confidence score and LPIPS perceptual sim- ilarity across each Top-n subset, ensuring the augmented data are both phenotype-consistent and visually plausible. Supervised scaling effect. To analyze the scaling effect of synthetic augmentation in standard supervised classifica- tion, models are retrained under a consistent protocol using (i) real data only (as baselines) and (i) real data combined with synthetic subsets of increasing size (1000-6000). Syn- thetic images are excluded from the test set to preserve eval- uation integrity. This setup enables a controlled analysis of how scaling phenotype-aligned synthetic data impacts clas- sification performance. Few-Shot learning. In few-shot learning, synthetic sam- ples are added to the support sets to increase training di- versity without altering the few-shot structure. For each class with m real support examples, (10âm) synthetic im- ages are sampled to reach 10 support examples per class. This augmentation strategy enables the exploration of dif- ferent support set sizes during training. Testing is per- formed strictly under the original n-way 1-shot configura- tion, using only real images for support and query sets to ensure unbiased evaluation. 5. Experiments and results This section presents the experimental evaluation of super- vised classification, few-shot learning, and synthetic data augmentation on the RDFace dataset. Detailed hyperpa- rameter settings and error bar calculations are provided in Appendix J and Appendix I respectively. 5.1. Standard supervised classification We first evaluated supervised classification across six back- bone architectures using real facial images only. As shown in Table 1, performance varied substantially across mod- els.DenseNet achieved the highest Top-1 accuracy at 15.93%, followed by Swin Transformer (14.34%) and VGG (11.68%). Other models, such as ResNet, FaceNet, and CLIP, trailed behind with lower accuracy, particularly CLIP (3.01%), which underperformed across all Top-k metrics. Table 1. Standard supervised classification results using real train- ing data across different backbones and baseline Gestalt. 1 Acc (%)Top-1Top-5Top-10Top-30 ResNet6.90 (1.45) 18.58 (3.00) 28.50 (3.67) 54.34 (2.39) DenseNet15.93 (2.34) 33.63 (3.70) 43.01 (2.63) 64.42 (1.92) FaceNet9.91 (1.81) 24.60 (5.43) 34.87 (4.75) 58.23 (5.06) VGG11.68 (1.58) 29.91 (2.68) 38.41 (1.34) 60.88 (2.02) Swin-T14.34 (2.61) 26.19 (2.04) 35.93 (3.17) 58.41 (3.81) CLIP 3.01 (1.48) 12.74 (2.84) 19.12 (5.18) 42.30 (4.40) Gestalt 6.19 (1.40) 17.52 (1.58) 27.79 (3.68) 49.20 (2.91) Top-5 and Top-10 accuracy results further reveal that model performance improves significantly when more guesses are permitted. For example, DenseNetâs Top-5 accuracy reached 33.63%, and Top-30 accuracy exceeded 64%. This suggests that while Top-1 classification remains 1 All results in this paper are reported as mean (standard deviation) un- less otherwise noted. The best results are highlighted in bold. difficult in ultra-low-shot settings, models can still learn useful representations for phenotype-based narrowing. 5.2. Few-shot learning Prototypical Networks were applied under 5, 10, and 15- way 1-shot settings (see Table 2).Few-shot learning achieved substantially higher accuracy within its simpli- fied episodic setup, highlighting its adaptability in ultra- low-shot settings. DenseNet achieved the highest accu- racy in the 5-way configuration (26.20%), while ResNet maintained more stable performance as task complexity in- creased (18.15% and 17.99% for 10- and 15-way, respec- tively). CLIP and FaceNet were less effective, especially in higher-way tasks, reflecting limited generalization of their embeddings in RD phenotyping. Table 2. Few-shot learning results under different settings using real training data across different backbone models. Acc (%)5-way 1-shot10-way 1-shot15-way 1-shot ResNet24.18 (2.56)18.15 (2.58)17.99 (1.26) DenseNet26.20 (2.01)17.36 (1.66)17.30 (3.45) FaceNet25.16 (4.89)12.93 (2.23)8.03 (1.40) VGG21.54 (3.72)10.82 (2.14)9.05 (2.20) Swin-T22.24 (2.88)10.94 (2.59)8.03 (2.41) CLIP23.48 (5.28)12.33 (1.96)7.30 (1.90) 5.3. Synthetic data generation and evaluation Representative synthetic images for Aarskog-Scott syn- drome (AAR) generated by DreamBooth, together with un- conditional samples from FastGAN, are presented in Fig- ure 4. Additional examples generated by both methods are provided in Appendix C. Overall, the generative pro- cess resulted in 10,300 DreamBooth-generated and 8,928 FastGAN-generated synthetic images. (a) DreamBooth(b) FastGAN Figure 4. Representative samples of synthetic images. (a) DreamBooth-generated synthetic images conditioned on AAR. (b) FastGAN-generated synthetic images. Beyond visual inspection, we evaluate the visual realism of generated samples. After filtering out samples flagged by safety checks or with RetinaFace confidence below 0.90, Figure 5. Comparison of phenotype descriptions generated by VLM between a real image and one corresponding synthetic image. Left: real image; Right: DreamBooth-generated synthetic image. The real image has been visually processed to reduce identifiability in accordance with privacy and ethical considerations. Green indicates consistent phenotype terms while red indicates conflicting descriptions. 99.42% of DreamBooth and 99.70% of FastGAN samples achieve high detection confidence (> 0.99). Perceptual sim- ilarity to real training images, measured using LPIPS, was 0.5337 for FastGAN and 0.4871 for DreamBooth. Phenotypic alignment was assessed via landmark-based cosine similarity and VLM-generated phenotype reports. DreamBooth-generated images achieved an average land- mark similarity rank of 19.74, confirming structural consis- tency with real disease prototypes. Additional details on the landmark similarity analysis can be found in Appendix D.1. Expert review of synthetic samples. To validate the clin- ical relevance of the synthetic images, a random subset of 50 DreamBooth and 50 FastGAN samples was evaluated. DreamBooth images were marked as Plausible in 62â76% of cases, depending on whether at least one or both review- ers agreed, while FastGAN images reached only 2â38%. Inter-rater reliability further supported these findings. For DreamBooth, the observed agreement was 84.0% with Co- henâs Îș = 0.65, indicating substantial agreement; while FastGAN showed only 38.0% observed agreement with Îș = 0.07, suggesting poor consistency. Detailed ratings are summarized in Appendix D.2. The higher plausibility of DreamBooth outputs suggests that conditioning helps pre- serve phenotype-specific features for clinical interpretation. VLM-cased phenotype alignment. The example in Fig- ure 5 shows consistent semantic alignment between real and synthetic descriptions, with most discrepancies occurring in less distinctive regions. More cases are provided in Ap- pendix G.2. We further evaluate phenotype alignment us- ing VLM-generated reports across different image pairs. As shown in Figure 6, real-syn similarity is comparable to real- real across both models, indicating that synthetic images preserve disease-specific phenotype information. Syn-syn similarity is consistently high, suggesting stable phenotype representations among generated samples. Together, these results support the utility of DreamBooth samples. Uncertainty and robustness. Via stochastic sampling, the stability of Qwen-generated reports yields a mean uncer- tainty score of 0.108± 0.017, indicating high consistency across samples. As shown in Appendix G.3, cross-model similarity remains consistent across real and synthetic im- ages, indicating that phenotype descriptions are not strongly dependent on the choice of VLM. Figure 6. Overall similarity scores for Qwen and LLaVA across different comparisons. Error bars indicate standard deviation. Potential regional bias. We further analyze potential bias across geographic regions on standard supervised learning and phenotype report similarity. Results indicate consistent trends across regions and are provided in the Appendix H. 5.4. Impact of synthetic data augmentation We evaluated the impact of synthetic data generated by DreamBooth and FastGAN across standard supervised and few-shot classification tasks under multiple configurations. For standard supervised classification, we first compare real only performances with generic augmentation methods as baselines. On our best backbone DenseNet: MixUp [53] and CutMix [52] achieve 15.75% and 16.11% Top-1 ac- curacy respectively, suggesting generic augmentations pro- vide limited or inconsistent gains under extreme low-shot conditions. DreamBooth (DB) augmentation (Table 3) con- sistently improved Top-1 accuracy across nearly all back- bones, demonstrating the benefit of phenotype-aligned syn- thesis. For example, DenseNet increased from 15.93% (real only) to 17.52% with DB augmentation, and VGG from 11.68% to 16.64%. In contrast, FastGAN (FG) augmen- tation led to performance degradation across most models. For instance, DenseNet dropped from 15.93% to 13.27%. The mixed (Real + DB + FG) configuration partially re- covered performance, reaching intermediate accuracy be- tween the two augmentations and being slightly better than the real-only baselines.These results show that class- conditioned augmentation is consistently beneficial, while unconditioned synthesis can distort class distributions. Table 3. Standard supervised classification results (Top-1 accura- cies) under landmark-based Top-1000 synthetic data augmentation across different backbone models. ACC (%)Real onlyReal+DBReal+FG Real+DB+FG ResNet6.90 (1.45) 12.21 (1.70) 8.50 (1.48)8.32(1.73) DenseNet 15.93 (2.34) 17.52 (2.29) 13.27 (3.13) 16.46 (2.63) FaceNet 9.91 (1.81) 15.04 (2.87) 6.55 (2.48)10.97 (2.55) VGG11.68 (1.58) 16.64 (4.07) 7.26 (2.37)12.92(2.84) Swin-T14.34 (2.61) 16.81 (1.40) 10.44 (1.70) 14.34 (1.70) CLIP3.01 (1.48) 9.03 (2.29) 1.42 (1.34)4.25 (1.58) Gestalt6.19 (1.40) 9.03 (1.45) 3.19 (1.84)5.31 (0.88) Ablation study on scaling effect. To analyze how the quantity and type of synthetic data influence model per- formance, we varied the number of DreamBooth- and FastGAN-generated samples added per class. As shown in Appendix F.1, both models exhibit distinct scaling be- haviors. DreamBooth shows a clear non-linear improve- ment pattern: accuracy rises sharply between the Top-1000 and Top-4000 subsets and then plateaus or slightly declines at Top-6000.For instance, DenseNetâs Top-1 accuracy increases from 15.93% (real only) to 17.52% (Top-1000) and reaches 21.06% at Top-6000, indicating that moder- ate quantities of phenotype-aligned samples most effec- tively enhance generalization. In contrast, FastGAN dis- plays an immediate and persistent decline as more sam- ples are added, reflecting low signal-to-noise in its uncon- ditioned generation. Overall, the ablation reveals that per- formance gains saturate with excessive synthetic data and are driven by phenotype fidelity rather than dataset volume, underscoring the value of phenotype-aware generation. Few-shot learning results (see Table 4 and Appendix F.2) further validate these observations. DreamBooth augmen- tation outperformed real-only baselines in most 1-shot set- tings. For instance, DenseNet improved from 26.20% to 29.88% (5-way 1-shot), and ResNet improved from 24.18% to 25.72%. Gains were also observed in 5-shot settings, where larger support sets enabled better utilization of syn- thetic diversity (e.g., ResNet reaching 33.62%). CLIP, how- ever, showed modest or inconsistent benefits, possibly due to its weaker initial alignment with facial phenotypes. Supplemental visual quality analysis (Appendix E) con- Table 4. Few-shot learning results under synthetic data augmenta- tion across different backbone models. 5-way 1-shot ACC (%)Real onlyReal + DB ResNet24.18 (2.56)25.72 (1.62) DenseNet 26.20 (2.01)29.88 (1.51) FaceNet25.16 (4.89)23.60 (4.06) VGG 21.54 (3.72)21.02 (5.85) Swin-T22.24 (2.88)26.72 (4.34) CLIP 23.48 (5.28)22.30 (2.63) firmed that Top-n landmark similarity ranking is predictive of image realism, which explains why DreamBooth samples generalize better across models and settings. 6. Conclusion Findings and implications. Overall, these findings high- light the versatility of the proposed rare disease dataset in supporting diverse analytical tasks, from classification and few-shot learning to phenotype evaluation. Supervised and few-shot classification results underscore the difficulty of ultra-low-shot rare disease diagnosis, with few-shot meth- ods offering modest gains in constrained settings. Synthetic data, particularly DreamBooth-generated samples, effec- tively mitigated data scarcity. Their consistent performance gains suggest improvements stem from higher phenotype fidelity rather than sample quantity or overfitting.The landmark-based similarity offers interpretable grounded validation and enable pseudo-labelling. Finally, the VLM- generated diagnostic reports demonstrated strong semantic consistency and interpretability, underscoring their poten- tial for future clinical and educational applications. Limitations. RDFace reflects the real-world scarcity and imbalance inherent in rare disease data and, while currently limited in scale, provides a valuable foundation for study- ing generalization under such constraints. The images were collected from heterogeneous online sources with varying completeness of demographic metadata, which may limit the extent of bias or cross-population generalizability. Conclusion. RDFace is a curated benchmark dataset de- signed for rare disease facial analysis under real-world sce- narios. It provides standardized settings for evaluating su- pervised, few-shot, and phenotype-aware learning methods, enabling systematic assessment of model performance un- der data scarcity. By integrating real pediatric images with synthetic augmentations and enabling both structural and semantic evaluations, RDFace supports comprehensive and reproducible assessment of model behavior. Our exper- iments demonstrate that phenotype-aligned synthetic data improves recognition while preserving clinically relevant features. Together, RDFace thus dserves as a foundation for developing reliable AI systems for rare disease diagnosis. Acknowledgement This work was supported in part by the Canada Research Chairs Tier I Program (CRC-2021-00482) and the Canada Foundation for Innovation John R. Evans Leaders Fund (JELF) program (#43481). All data collection procedures were approved by the Western University Health Science Research Ethics Board (HSREB) (Reference No. 2023- 122744-77394). Facial photographs of children with rare diseases were collected from publicly available sources, in- cluding the published literature and foundation websites, and the authors gratefully acknowledge these sources. The authors sincerely thank Dr. Patrick Frosk for his support in the design of the project. The authors also acknowl- edge the Digital Research Alliance of Canada and Compute Canada for providing the computational resources used in this study. References [1] Matilda Anderson, Elizabeth J. Elliott, and Yvonne A. Zurynski. Australian families living with rare disease: expe- riences of diagnosis, health services use and needs for psy- chosocial support. Orphanet Journal of Rare Diseases, 8:22, 2013. 1 [2] Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhao- hai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Jun- yang Lin. Qwen2.5-VL Technical Report. arXiv e-prints, art. arXiv:2502.13923, 2025. 5 [3] Mark L Batshaw, Stephen C Groft, and Jeffrey P Krischer. Research into rare diseases of childhood. JAMA : the journal of the American Medical Association, 311(17):1729â1730, 2014. 3 [4] Gareth Baynam, Nicholas Pachter, Fiona McKenzie, Sharon Townshend, Jennie Slee, Cathy Kiraly-Borri, Anand Va- sudevan, Anne Hawkins, Stephanie Broley, Lyn Schofield, Hedwig Verhoef, Caroline E. Walker, Caron Molster, Jene- fer M. Blackwell, Sarra Jamieson, Dave Tang, Timo Lass- mann, Kym Mina, John Beilby, Mark Davis, Nigel Laing, Lesley Murphy, Tarun Weeramanthri, Hugh Dawkins, and Jack Goldblatt. The rare and undiagnosed diseases diagnos- tic service â application of massively parallel sequencing in a state-wide clinical service. Orphanet Journal of Rare Dis- eases, 11:77, 2016. 1 [5] Qiong Cao, Li Shen, Weidi Xie, Omkar M. Parkhi, and An- drew Zisserman.VGGFace2: A Dataset for Recognising Faces across Pose and Age . In 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), pages 67â74, Los Alamitos, CA, USA, 2018. IEEE Computer Society. 3 [6] Raquel Castro, Juliette Senecat, Myriam de Chalendar, Ildik Ì o Vajda, Dorica Dan, and B Ì eata Boncz. Bridging the Gap between Health and Social Care for Rare Diseases: Key Issues and Innovative Solutions, pages 605â627. Springer International Publishing, Cham, 2017. 1 [7] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248â255, 2009. 3 [8] Jiankang Deng, Jia Guo, Evangelos Ververas, Irene Kotsia, and Stefanos Zafeiriou. Retinaface: Single-shot multi-level face localisation in the wild. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5202â5211, 2020. 4 [9] EURORDIS. The voice of 12,000 patients. experiences and expectations of rare disease patients on diagnosis and care in europe, 2009. 1 [10] Carole Faviez, Xiaoyi Chen, Nicolas Garcelon, Antoine Neuraz, Bertrand Knebelmann, R Ì emi Salomon, Stanislas Ly- onnet, Sophie Saunier, and Anita Burgun. Diagnosis support systems for rare diseases: A scoping review. Orphanet Jour- nal of Rare Diseases, 15(1):94, 2020. 1 [11] Maciej Geremek and Krzysztof Szklanny. Deep learning- based analysis of face images as a screening tool for genetic syndromes. Sensors, 21(19), 2021. 2 [12] Danila Germanese, Sara Colantonio, Marco Del Coco, Pier- luigi Carcagn ` ı, and Marco Leo. Computer vision tasks for ambient intelligence in childrenâs health. Information, 14 (10), 2023. 2 [13] Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Proceed- ings of the 28th International Conference on Neural Infor- mation Processing Systems - Volume 2, page 2672â2680, Cambridge, MA, USA, 2014. MIT Press. 2 [14] Yaron Gurovich, Yair Hanani, Omri Bar, Guy Nadav, Nicole Fleischer, Dekel Gelbman, Lina Basel-Salmon, Peter M. Krawitz, Susanne B. Kamphausen, Martin Zenker, Lynne M. Bird, and Karen W. Gripp. Identifying facial phenotypes of genetic disorders using deep learning. Nature Medicine, 25: 60 â 64, 2019. 1 [15] Melissa Haendel, Nicole Vasilevsky, Deepak Unni, Cristian Bologa, Nomi Harris, Heidi Rehm, Ada Hamosh, Gareth Baynam, Tudor Groza, Julie McMurry, Hugh Dawkins, Ana Rath, Courtney Thaxton, Giovanni Bocci, Marcin P. Joachimiak, Sebastian K Ì ohler, Peter N. Robinson, Chris Mungall, and Tudor I. Oprea. How many rare diseases are there? Nature Reviews Drug Discovery, 19(2):77â78, 2020. 1 [16] Charles R. Harris, K. Jarrod Millman, St Ì efan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fern Ì andez del R Ì Ä±o, Mark Wiebe, Pearu Peterson, Pierre G Ì erard- Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant. Array programming with NumPy. Nature, 585(7825):357â 362, 2020. 20 [17] Iryna Hartsock and Ghulam Rasool. Vision-language mod- els for medical report generation and visual question answer- ing: a review. Frontiers in Artificial Intelligence, Volume 7 - 2024, 2024. 2 [18] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770â778, 2016. 3 [19] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. In Advances in Neural Infor- mation Processing Systems, pages 6840â6851. Curran Asso- ciates, Inc., 2020. 2 [20] Tzung-Chien Hsieh, Aviram Bar-Haim, Shahida Moosa, Nadja Ehmke, Karen W. Gripp, Jean Tori Pantel, Mag- dalena Danyel, Martin Atta Mensah, Denise Horn, Stanislav Rosnev, Nicole Fleischer, Guilherme Bonini, Alexander Hustinx, Alexander Schmid, Alexej Knaus, Behnam Ja- vanmardi, Hannah Klinkhammer, Hellen Lesmann, Su- girthan Sivalingam, Tom Kamphans, Wolfgang Meiswinkel, Fr Ì ed Ì eric Ebstein, Elke Kr Ì uger, S Ì ebastien K Ì ury, St Ì ephane B Ì ezieau, Axel Schmidt, Sophia Peters, Hartmut Engels, Elis- abeth Mangold, Martina KreiĂ, Kirsten Cremer, Claudia Perne, Regina C. Betz, Tim Bender, Kathrin Grundmann- Hauser, Tobias B. Haack, Matias Wagner, Theresa Brunet, Heidi Beate Bentzen, Luisa Averdunk, Kimberly Christine Coetzer, Gholson J. Lyon, Malte Spielmann, Christian P. Schaaf, Stefan Mundlos, Markus M. N Ì othen, and Peter M. Krawitz. Gestaltmatcher facilitates rare disease matching us- ing facial phenotype descriptors. Nature Genetics, 54:349â 357, 2022. 2 [21] Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kil- ian Q. Weinberger. Densely connected convolutional net- works. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2261â2269, 2017. 3 [22] Showrov Islam, Md. Tarek Aziz, Hadiur Rahman Nabil, Jamin Rahman Jim, M. F. Mridha, Md. Mohsin Kabir, Nobuyoshi Asai, and Jungpil Shin. Generative adversar- ial networks (gans) in medical imaging: Advancements, ap- plications, and challenges. IEEE Access, 12:35728â35753, 2024. 2 [23] Bo Jin, Leandro Cruz, and Nuno Gonc ̧alves. Deep facial diagnosis: Deep transfer learning from face recognition to facial diagnosis. IEEE Access, 8:123649â123661, 2020. 2 [24] Xiaoyang Kang, Tao Yang, Wenqi Ouyang, Peiran Ren, Lingzhi Li, and Xuansong Xie. Ddcolor: Towards photo- realistic image colorization via dual decoders.In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 328â338, 2023. 4 [25] Aron Kirchhoff, Alexander Hustinx, Behnam Javanmardi, Tzung-Chien Hsieh, Fabian Brand, Fabio Hellmann, Silvan Mertes, Elisabeth Andr Ì e, Shahida Moosa, Thomas Schultz, Benjamin D. Solomon, and Peter Krawitz. Gestaltgan: syn- thetic photorealistic portraits of individuals with rare genetic disorders. European Journal of Human Genetics, 33:377â 382, 2025. 2 [26] Maya Koretzky, Vence L. Bonham, Benjamin E. Berkman, Paul Kruszka, Adebowale Adeyemo, Maximilian Muenke, and Sara Chandros Hull. Towards a more representative mor- phology: clinical and ethical considerations for including di- verse populations in diagnostic genetic atlases. Genetics in Medicine, 18(11):1069â1074, 2016. 2 [27] Peter Kov Ì a Ë c, Peter Jackuliak, Alexandra Bra Ë zinov Ì a, Ivan Varga, Michal Al Ì a Ë c, Martin Smatana, Du Ë san Lovich, and Andrej Thurzo. Artificial intelligence-driven facial image analysis for the early detection of rare diseases: Legal, ethi- cal, forensic, and cybersecurity considerations. AI, 5(3):990â 1010, 2024. 1 [28] Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234â1240, 2019. 5 [29] Bingchen Liu, Yizhe Zhu, Kunpeng Song, and A. Elgammal. Towards faster and stabilized GAN training for high-fidelity few-shot image synthesis. In International Conference on Learning Representations (ICLR), 2021. 2 [30] Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024. 5 [31] Ying Liu, Hengchang Zhang, Weidong Zhang, Guojun Lu, Qi Tian, and Nam Ling. Few-shot image classification: Cur- rent status and research trends. Electronics, 11(11), 2022. 2 [32] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows . In 2021 IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 9992â10002, Los Alamitos, CA, USA, 2021. IEEE Computer Society. 3 [33] Christopher D Manning. An introduction to information re- trieval. Cambridge University Press, 2009. 18 [34] Jannatul Nayem, Sayed Sahriar Hasan, Noshin Amina, Bristy Das, Md Shahin Ali, Md Manjurul Ahsan, and Shiv- akumar Raman. Few Shot Learning for Medical Imaging: A Comparative Analysis of Methodologies and Formal Mathe- matical Framework, pages 69â90. Springer Nature Switzer- land, Cham, 2023. 2 [35] Omkar M. Parkhi, Andrea Vedaldi, and Andrew Zisserman. Deep face recognition. In Proceedings of the British Machine Vision Conference (BMVC), pages 41.1â41.12. BMVA Press, 2015. 3 [36] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K Ì opf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: an imper- ative style, high-performance deep learning library. Curran Associates Inc., Red Hook, NY, USA, 2019. 21 [37] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825â2830, 2011. 18 [38] Jiaqi Qiang, Danning Wu, Hanze Du, Huijuan Zhu, Shi Chen, and Hui Pan. Review on facial-recognition-based ap- plications in disease diagnosis. Bioengineering, 9(7), 2022. 1 [39] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, pages 8748â8763. PMLR, 2021. 3 [40] Ruiyang Ren, Haozhe Luo, Chongying Su, Yang Yao, and Wen Liao. Machine learning in dental, oral and craniofacial imaging: a review of recent progress. PeerJ, 9:e11451, 2021. 2 [41] Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In 2023 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pages 22500â22510, 2023. 2 [42] Arrigo Schieppati, Jan-Inge Henter, Erica Daina, and Anita Aperia. Why rare diseases are an important medical and so- cial issue. The Lancet, 371(9629):2039â2041, 2008. 1 [43] Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clus- tering. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 815â823, 2015. 3 [44] Fayroz F. Sherif, Nahed Tawfik, Doaa Mousa, Mohamed S. Abdallah, and Young-Im Cho. Automated multi-class facial syndrome classification using transfer learning techniques. Bioengineering, 11(8), 2024. 2 [45] Karen Simonyan and Andrew Zisserman. Very deep con- volutional networks for large-scale image recognition. In 3rd International Conference on Learning Representations, ICLR 2015, 2015. 3 [46] Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. In Advances in Neural Infor- mation Processing Systems. Curran Associates, Inc., 2017. 2 [47] Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-esrgan: Training real-world blind super-resolution with pure synthetic data. In 2021 IEEE/CVF International Con- ference on Computer Vision Workshops (ICCVW), pages 1905â1914, 2021. 4 [48] S S Weinreich, R Mangon, J J Sikkens, M E en Teeuw, and M C Cornel. Orphanet: a european database for rare dis- eases. Nederlands tijdschrift voor geneeskunde, 152(9):518â 519, 2008. 3 [49] Hang Yang, Xin-Rong Hu, Ling Sun, Dian Hong, Ying-Yi Zheng, Ying Xin, Hui Liu, Min-Yin Lin, Long Wen, Dong- Po Liang, and Shu-Shui Wang. Automated facial recognition for noonan syndrome using novel deep convolutional neural network with additive angular margin loss. Frontiers in Ge- netics, 12:669841, 2021. 2 [50] Nur Yildirim,Hannah Richardson,Maria Teodora Wetscherek, Junaid Bajwa, Joseph Jacob, Mark Ames Pinnock, Stephen Harris, Daniel Coelho De Castro, Shruthi Bannur, Stephanie Hyland, Pratik Ghosh, Mercy Ranjit, Kenza Bouzid, Anton Schwaighofer, Fernando P Ì erez- Garc Ì Ä±a, Harshita Sharma, Ozan Oktay, Matthew Lungren, Javier Alvarez-Valle, Aditya Nori, and Anja Thieme. Multimodal healthcare ai: Identifying and designing clin- ically relevant vision-language applications for radiology. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, New York, NY, USA, 2024. Association for Computing Machinery. 2 [51] Haojun Yu, Youcheng Li, Nan Zhang, Zihan Niu, Xuan- tong Gong, Yanwen Luo, Haotian Ye, Siyu He, Quanlin Wu, Wangyan Qin, Mengyuan Zhou, Jie Han, Jia Tao, Zi- wei Zhao, Di Dai, Di He, Dong Wang, Binghui Tang, Ling Huo, James Zou, Qingli Zhu, Yong Wang, and Liwei Wang. A Foundational Generative Model for Breast Ultrasound Im- age Analysis. arXiv e-prints, art. arXiv:2501.06869, 2025. 4 [52] Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regu- larization strategy to train strong classifiers with localizable features. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 6022â6031, 2019. 7 [53] Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimiza- tion. In International Conference on Learning Representa- tions, 2018. 7 [54] Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 586â595, 2018. 5 [55] Yvonne Zurynski, Aranzazu Gonzalez, Marie Deverell, Amy Phu, Helen Leonard, John Christodoulou, and Elizabeth El- liott. Rare disease: a national survey of paediatriciansâ ex- periences and needs. BMJ Paediatrics Open, 1(1):e000172, 2017. 1 RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation Supplementary Material Appendix Contents A. Dataset documentation2 A.1. Disease list and metadata . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .2 A.2. Dataset distribution and organization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .5 B. Few-shot learning6 B.1. Prototypical networks algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .6 C. Synthetic image samples7 D. Expert and automated evaluation of synthetic data9 D.1. Landmark-based similarity analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .9 D.1.1. Heatmap of landmark-based cosine similarities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .9 D.1.2. Ranking consistency of DreamBooth-generated Images . . . . . . . . . . . . . . . . . . . . . . . . . 10 D.2. Expert review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 D.3. Observations and implications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 E. Tradeoff between disease-specific structure and visual realism12 F. Synthetic data-involved downstream tasks results13 F.1. Standard supervised classification and synthetic scaling effect . . . . . . . . . . . . . . . . . . . . . . . . . . 13 F.2. Few-shot learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 F.3. Observations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 G. VLM-based report generation16 G.1. Prompt design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 G.2. Report evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 G.3. Uncertainty and robustness analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 H. Regional analysis and potential bias19 I . Statistical reporting details20 J. Hyperparameter settings and training details20 J.1 . Standard supervised classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 J.2 . Few-Shot learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 J.3 . Synthetic data generation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 J.4 . Hardware and compute resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 A. Dataset documentation A.1. Disease list and metadata We provide a complete list of the 103 rare disease classes included in the RDFace dataset. As summarized in Table S1, each entry includes the disease name, abbreviation (Abbr) (used for labeling), associated gene (if available), clinical subcategory based on Orphanet (if available), Orphanet code, and the number of real facial images (# Img) curated for that class. Table S1. Metadata of rare disease classes in the RDFace dataset. 2 Disease NameAbbrGeneSubcategoryOrphanet Code# Img Aarskog-Scott syndromeAARFGD1Delayed puberty9155 Aicardi-GoutieresAICTREX1?515 Triple A syndromeALLAAASmultisystem disease8695 Allan-herndon-dudley SyndromeALLASLC16A2Ataxia595 Alport syndromeALP?Abnormal retinal morphology634 Alpha-thalassemiaALPH?Cholestasis8465 Alpha-mannosidase deficiencyALPMMAN2B1lysosomal storage disease615 AlstromALSALSM1multisystemic disorder645 Angelman SyndromeANGUBE3AIntellectual disability725 X-linked cleft palate and anky- loglossia ANKTBX22developmental defect during embryogenesis syndrome 3246012 Apert syndromeAPEFGFR2Cardiomyopathy875 AR Polycystic Kidney DiseaseARPPKHD1hepatorenal fibrocystic syn- drome 7311 Arterial TortuosityART?connective tissue disorder33425 Ataxia TelangiectasiaATAATMcombined dystonia1002 Atypical Rett SyndromeATYPMECP2, GABBR2, STXBP1, CDKL5 Calcium nephrolithiasis30955 Auriculocondylar SyndromeAURIEDN1, PLCB4, GNAI3 Abnormal soft palate mor- phology to abnormality of the uvulva 1378885 Bainbridge-ropers SyndromeBAIASXL3Severe postnatal growth retar- dation 3525775 Bardet-BiedlBARBBS2ciliopathy with multisystem involvement 1105 Barber-say SyndromeBARBTWIST2Hyperextensible skin12315 Barth syndromeBART?Cardiomyopathy1115 Beckwith-wiedemann SyndromeBEC?Multiple renal cysts1165 Bohring-opitz SyndromeBOHASXL1Muscular hypotonia972975 Boycott-Beaulieu-InnesBOYTHOC6syndromic intellectual disabil- ity disorder 3634445 Congenital adrenal hyperplasiaCAH?Congenital adrenal hyperpla- sia 4185 Canavan DiseaseCANASPAAbnormality of serum amino acid level 1415 COFS syndromeCERERCC6diseases of DNA repair14665 Chudley-McCulloughCHUGPSM2syndromic deafness3145974 Cleidocranial DysplasiaCLECBFA1developmental abnormality of bone 14525 Clouston SyndromeCLOUGJB6Ectodermal dysplasia1894 CODASCODLONP1multiple congenital anomalies syndrome 14585 Coffin-lowry SyndromeCOFRPS6KA3Spasticity1925 Combined pituitary hormone defi- ciencies, genetic forms COMPROP1Congenital hypopituitarism954944 SLC39A8-CDGCONDSLC39A8?4686995 Disease NameAbbrGeneSubcategoryOrphanet Code# Img Cranioectodermal dysplasiaCRADPH1developmental disorder15155 Crisponi SyndromeCRISCLCF1, CRLF1Malignant hyperthermia15455 3C syndromeCSYCCDC22, WASHC5 ?75 Cushing diseaseCUSUSP8, CDH23Abnormal bleeding962534 Cystic FibrosisCYSCFTR?5865 Diamond-blackfan AnemiaDIARPS19Colon cancer1245 Donnai-barrow SyndromeDONLRP2Partial agenesis of the corpus callosum 21435 DOORS syndromeDOOTBC1D24Hypothyroidism795005 Dopa-Responsive DystoniaDOPTHgroup of diseases2553 DysosteosclerosisDYS?primary bone dysplasia dis- ease 17822 Earlyinfantileepilepticen- cephalopathy EPIE?epileptic encephalopathy19345 Fanconi AnemiaFANCFANCCDNA repair disorder845 Geroderma OsteodysplasticaGEOGORAB?20785 Glutaryl-CoA dehydrogenase defi- ciency GLUGCDHneurometabolic disorder255 H(hyperornithinemia- hyperammonemia- homocitrullinuria) HHHSLC25A15disorderofureacycle metabolism 4151 Hypohidrotic Ectodermal Dyspla- sia HYPEEDA1disorder of ectoderm develop- ment 2384681 HypophosphatasiaHYPOALPLmetabolic disorder4365 Severecombinedimmunodefi- ciency IMMUDCLRE1Cprimary immunodeficiency1836605 Joubert SyndromeJOUATMEM237?4755 Juvenile Amyotrophic Lateral Scle- rosis JUVALS2, HNRNPA2B1, HNRNPA1 Motor neuron atrophy3006055 Kawasaki DiseaseKAW?Arrhythmia23315 Infantile Krabbe DiseaseKRAGALC, PSAPGeneralizedmyoclonic seizure 2064365 Laron SyndromeLARGHRHypoglycemia6333 LeighLEINDUFV1progressive neurological dis- ease 5065 LeprechaunismLEPINSRHypoglycemia5087 Loeys-DietzLOETGFB2connective tissue disorder600305 Malignant hyperthermia of anesthe- sia MALRYR1pharmacogenetic disorder of skeletal muscle 4233 Maple Syrup Urine DiseaseMAPBCKDHBdisorder of branched-chain amino acid metabolism 5115 Marden-walker SyndromeMARPIEZO2Muscular dystrophy24615 Marinesco-sj Ì ogren SyndromeMARIINPP5K, SIL1Muscular dystrophy5595 Oculotrichoanal syndromeMBOFREM1multiple congenital anomalies27175 Mucopolysaccharidosis type 4MOR?lysosomal storage disease5825 Mowat-wilson SyndromeMOWZEB2Abdominal distention21525 Mucopolysaccharidosis type 7MPSGUSBlysosomal storage disease5845 Multiple Sulfatase DeficiencyMULSUMF1Progressive neurologic deteri- oration 5855 Ochoa SyndromeOCHHPSE2, LRIG2Renal insufficiency27045 Oculocutaneous AlbinismOCUTYRdisorders of melanin biosyn- thesis 555 Odonto-onycho-dermal dysplasiaODOWNT10Aectodermal dysplasia27215 Osteogenesis imperfectaOSTSEPINF1group of diseases6665 Parietal ForaminaPAR??600153 Disease NameAbbrGeneSubcategoryOrphanet Code# Img Glycogen storage disease due to acid maltase deficiency, late-onset POMGAAGlycogen storage disease4204295 Pontocerebellar HypoplasiaPONTOE1?985235 Primary HyperoxaluriaPRIGRHPRdisorderofglyoxylate metabolism 4162 Microcephaly-lymphedema- chorioretinopathy syndrome PRIM??25265 Propionic AcidemiaPROPCCBorganic aciduria355 PRUNE1-related neurological syn- drome PRUPRUNE1?5444692 Pten Hamartoma Tumor SyndromePTEN?Endometrial carcinoma3064985 Pyruvate Carboxylase DeficiencyPYRPCneurometabolic disorder30082 RenpenningRENPQBP1intellectualdisabilitysyn- drome 32424 Restrictive DermopathyRESZMPSTE24congenital genodermatosis16625 RhizomelicChondrodysplasia Punctata RHOPEX7group of diseases1775 RobertsROBESCO2?31035 Spinalarteriovenousmetameric syndrome SAMGSC?537211 Sandhoff DiseaseSAND?Progressive psychomotor de- terioration 7965 Sialidosis type 2SIANEU1lysosomal storage disease878765 Sickle Cell AnemiaSICHBBChronic hemolytic anemia2323 Proximal Spinal Muscular AtrophySPISMN1neuromuscular disorder704 Spondylodysplastic Ehlers-danlos Syndrome SPO?Platyspondyly5364713 Congenital sucrase-isomaltase defi- ciency SUCSIcarbohydrate intolerance dis- order 351222 Isolated sulfite oxidase deficiencySULSO?997315 Temple syndromeTEMP?Maturity-onset diabetes of the young 2545165 Turner SyndromeTUR?Biliary cirrhosis8814 Tyrosinemia Type 1TYRFAHinbornerroroftyrosine catabolism 8825 Usher Syndrome type 1USHBMYO7A?2311694 CACH syndromeVANEIF2B5?1355 Hyperostosis corticalis generalisataVANB?craniotubular hyperostosis34162 Walker WarburgWALPOMT1congenital muscular dystro- phy 8995 Warsaw Breakage SyndromeWARDDX11psychomotor retardation2805585 Wolf-hirschhorn SyndromeWOLLETM1, NSD2Rib segmentation abnormali- ties 2805 ZellwegerZEL?peroxisome biogenesis disor- der 9125 2 ? indicates that the information is not available. A.2. Dataset distribution and organization Aside from the metadata table, we also provide two figures to illustrate the distribution and structure of the RDFace dataset. Figure S1a shows the number of images available for each disease class, highlighting the inherent imbalance in data availabil- ity across different rare diseases. Figure S1b depicts the organization of the dataset. These visualizations provide a clearer understanding of the datasetâs composition and the challenges posed by data scarcity in rare disease research. (a) Distribution of images across disease classes.(b) Structure of the RDFace dataset. Figure S1. Overview of the RDFace dataset. B. Few-shot learning Few-shot learning (FSL) addresses the problem of classifying instances when only a few labeled examples are available for each class. This setting is common in domains like rare disease diagnosis, where collecting large-scale labeled datasets is impractical due to the low prevalence and data privacy constraints. In the standard n-way k-shot FSL setting, each learning episode involves n distinct classes, with only k labeled examples (support set) per class. The goal is to classify unlabeled examples (query set) drawn from the same n classes, using the limited support data for reference. Formally, each episode consists of: âą A support setS =(x i ,y i ) n·k i=1 , containing k samples per class. âą A query setQ =(x j ,y j ) q j=1 , containing unseen samples from the same n classes. We adopted Prototypical Networks to solve the few-shot classification task on RDFace. This approach learns an embed- ding function f Ξ :X âR d that maps input images into a feature space. For each class c i , a prototype (centroid) is computed as the mean embedding of its support samples: ÎŒ i = 1 k X xâS i f Ξ (x) (1) For each query image x (q) , the model computes the distance between its embedding and each prototype ÎŒ i , typically using the squared Euclidean distance: d(x (q) ,ÎŒ i ) = f Ξ (x (q) )â ÎŒ i 2 2 (2) The query image is then assigned to the class with the closest prototype, and a cross-entropy loss is computed over the predictions for optimization. B.1. Prototypical networks algorithm We provide the algorithm used for episodic training on RDFace as below: Algorithm 1 Prototypical Networks Episodic Training on RDFace Require: Feature extractor f Ξ , training classesC train , number of ways n, support size k = 1 1: Sample n classesc 1 ,...,c n âŒC train 2: for each class c i do 3:Sample support setS i =x (s) i , query setQ i =x (q) i 4:Compute prototype: ÎŒ i â f Ξ (x (s) i ) 5: end for 6: for each query x (q) j â S i Q i do 7:Compute distances: d ij ââ„f Ξ (x (q) j )â ÎŒ i â„ 2 8:Predict label: Ëy j â arg min i d ij 9: end for 10: Compute loss: L episode â CrossEntropy(Ëy,y) 11: Update f Ξ via backpropagation C. Synthetic image samples Figure S2 presents the top 100 synthetic facial images generated by FastGAN, ranked by their similarity to real disease faces using a landmark-based cosine metric. These samples highlight the modelâs ability to capture coarse facial structure across diverse rare disease classes, although occasional artifacts and inconsistencies remain visible. In contrast, Figure S3 provides a more targeted comparison using DreamBooth of 50 classes choosen to display. For each disease class, we show the most and least phenotype-consistent synthetic images compared to real images based on 5Ă 5 facial landmark cosine similarity. This comparison illustrates both the strength and the limitations of current generative models. While top-ranked DreamBooth samples often replicate key craniofacial traits, the bottom-ranked samples reveal potential risks of phenotype distortion or mode collapse. These visualizations support the use of generative models for phenotype augmentation, while also emphasizing the need for careful validation when applying synthetic images in clinical or diagnostic settings. Figure S2. Representative FastGAN images ranked most similar to diseases. Figure S3. Top-1 and Bottom-1 synthetic images based on similarity to real images. Top-1 and Bottom-1 denote the most and least similar samples respectively. D. Expert and automated evaluation of synthetic data D.1. Landmark-based similarity analysis We conducted our evaluation across all 103 rare disease classes in RDFace. For each class, DreamBooth-generated synthetic images were used, provided they passed quality checks based on facial structure and alignment. Images failing these checks, such as those with low RetinaFace detection confidence or invalid landmark configurations, were excluded at the image level, not the class level. All similarity analyses were performed using normalized 5Ă 5 landmark distance matrices computed from RetinaFace keypoints. D.1.1. Heatmap of landmark-based cosine similarities Figure S4 presents a cosine similarity heatmap comparing the average DreamBooth-generated landmark structure for each class against all real disease class prototypes. The heatmap reveals a strong diagonal trend, reflecting high intra-class similarity, with several classes showing close alignment between synthetic and real facial geometry. However, for certain classesâparticularly those with very limited or noisy training dataâthe diagonal intensity diminishes, indicating lower phe- notype fidelity. Off-diagonal activity highlights cross-class resemblance, often reflecting overlapping craniofacial features among syndromes or mode collapse in synthesis. Figure S4. Cosine similarity heatmap between DreamBooth-generated and real disease prototypes. The heatmap compares the average landmark distance matrices of DreamBooth-generated images (rows) to those of real disease class prototypes (columns). Brighter values indicate greater similarity. D.1.2. Ranking consistency of DreamBooth-generated Images To quantify how consistently DreamBooth preserves class-specific features, we computed the rank position of the ground- truth class for each synthetic image among all real class prototypes based on cosine similarity. Figure S5 summarizes the average rank per class. While many classes exhibit ranks close to 1, several fall well below the mean, suggesting inconsistency in capturing distinctive morphology. These cases often correspond to underrepresented or visually subtle conditions in the training set. The overall mean rank across all classes is 19.74, serving as a practical reference for DreamBoothâs phenotype alignment under ultra-low-shot conditions. Figure S5. Average rank of the true disease class for each DreamBooth-generated image. Each rank is calculated using cosine similarity to the corresponding real class prototype. The red dashed line indicates the overall mean rank. D.2. Expert review In order to evaluate the clinical plausibility of synthetic image-label pairs, we conducted an expert review of 50 DreamBooth and 50 FastGAN samples. Each image was independently assessed by two medical doctor (MD) students (MD1 and MD2) and categorized as Plausible, Implausible, or Uncertain. The results of this expert evaluation are summarized in Table S2. Table S2. Expert evaluation of the plausibility of synthetic image-label pairs generated by DreamBooth and FastGAN with two MDs. Plausibility RatingDreamBoothFastGAN MD1MD2MD1MD2 Plausible3831145 Implausible231530 Uncertain 10162115 To assess the consistency between the two raters, inter-rater reliability was computed for each model, and the pairwise confusion matrices are presented in Table S3. Each cell represents the number of images assigned to a given label combination by the two MD reviewers (DreamBooth / FastGAN). A strong concentration of counts along the diagonal indicates higher agreement on image plausibility. Quantitatively, the inter-rater agreement results (Table S4) demonstrate that DreamBooth achieved a substantially higher degree of consensus between raters ( Îș = 0.654), corresponding to âsubstantial agreementâ on the LandisâKoch scale, whereas FastGAN showed only minimal agreement (Îș = 0.069). This difference highlights that DreamBooth-generated Table S3. Expert validation confusion matrices for DreamBooth (DB) and FastGAN (FG). (DB / FG) LabelPlausibleImplausibleUncertainTotal Plausible31 / 20 / 00 / 331 / 5 Implausible0 / 72 / 111 / 123 / 30 Uncertain7 / 50 / 49 / 616 / 15 Total38 / 142 / 1510 / 2150 / 50 samples were generally perceived as more clinically plausible and consistent across evaluators, while FastGAN outputs ex- hibited greater variability and uncertainty. Table S4. Inter-rater agreement metrics for DreamBooth (DB) and FastGAN (FG). MethodObserved Agreement (%)Cohenâs ÎșSE95% CI DB84.0 (42/50)0.6540.106[0.446, 0.862] FG38.0 (19/50)0.0690.091[â0.110, 0.248] Overall, the expert review confirms that DreamBooth produces synthetic facial images that retain phenotypic plausibility and diagnostic relevance more consistently than FastGAN, aligning with the quantitative similarity metrics and qualitative visual inspection presented in the main manuscript. D.3. Observations and implications Our evaluation highlights the complementary strengths of automated and expert-based assessments for characterizing the quality of synthetic facial images. Landmark-based similarity analysis reveals that many disease classes exhibit distinct and consistent craniofacial geometry, validating the use of facial landmarks as a phenotypic signature. The observed variability in average rank and cosine similarity across classes reflects sensitivity to both morphological subtlety and the quality of training exemplars. These findings indicate that normalized landmark distance metrics offer an interpretable, spatially grounded method for quantifying phenotype preservation and filtering synthetic samples prior to downstream tasks. However, structural alignment alone does not guarantee clinical plausibility. To address this, we conducted an expert re- view to evaluate whether synthetic image-label pairs appeared medically credible. Results show that DreamBooth-generated images were more frequently judged as plausible by at least one medical doctor, whereas FastGAN samples showed higher rates of uncertainty and implausibility. These findings underscore that perceptual realism, which is often emphasized in generative model benchmarks, does not always correlate with diagnostic fidelity. Expert review provides a critical semantic layer that captures domain-specific context often missed by automated metrics. Together, these results suggest that effective synthetic data curation requires a multi-faceted evaluation strategy that in- tegrates spatial similarity, visual quality, and human judgment. Such approaches are especially important in rare disease settings, where subtle phenotypic cues and label noise can dramatically impact model performance. In future applications, combining interpretable landmark filtering with lightweight expert-in-the-loop review may provide a scalable path for en- hancing dataset quality and clinical utility. E. Tradeoff between disease-specific structure and visual realism To assess the quality and disease-relevance of generated synthetic images, we evaluate two image-level realism metrics: RetinaFace detection confidence and LPIPS perceptual similarity. These metrics are computed across Top-n subsets (ranging from 1000 to 6000 images), where the images are ranked by their landmark-based cosine similarity to real disease class prototypes. DreamBooth Figure S6 shows the trend of RetinaFace and LPIPS scores across Top-n DreamBooth subsets. As n in- creases, RetinaFace confidence scores gradually decline, suggesting reduced alignment and detectability in lower-ranked images. Meanwhile, LPIPS scores increase, indicating a decrease in perceptual similarity to real images. These opposing trends reflect a trade-off between structural fidelity (as captured by landmarks) and low-level texture realism. The landmark- based ranking effectively promotes DreamBooth samples with coherent facial geometry and higher visual plausibility. Figure S6. DreamBooth â Correlation Between Top-n Ranking and Visual Realism. RetinaFace detection confidence (left) and LPIPS similarity (right) across Top-n ranked DreamBooth images. FastGAN FastGAN samples exhibit a broadly similar trend to DreamBooth in terms of ranking-based visual quality (see Figure S7). As the Top-n threshold increases, RetinaFace confidence scores slightly decline and LPIPS scores gradually increase, indicating reduced structural detectability and perceptual similarity at lower-ranked positions. However, some local fluctuations are observedâparticularly around the Top-2000 cutoffâwhere both metrics deviate slightly from the overall trajectory. These irregularities suggest that landmark-based ranking is still somewhat effective for FastGAN, but may be less stable than for DreamBooth due to the lack of class-specific conditioning. Figure S7. FastGAN â Correlation Between Top-n Ranking and Visual Realism. RetinaFace detection confidence (left) and LPIPS similarity (right) across Top-n ranked FastGAN images. F. Synthetic data-involved downstream tasks results F.1. Standard supervised classification and synthetic scaling effect Table S5 and Table S6 below report Top-k classification accuracies across various backbones. Each row represents classifica- tion performance (Top-k accuracy) under a different training configuration. âReal onlyâ refers to models trained exclusively on real RDFace data. âTop-nâ rows correspond to training sets augmented with the Top-n synthetic images selected based on landmark similarity to real samples. The synthetic images are chosen to best align with phenotype-specific facial structure. Table S5. Top-k accuracies (%) across backbones and synthetic cutoffs of DreamBooth samples. ACC (%)Top-nResNetDenseNetFaceNetVGGSwin-TCLIP Top-1 Real only6.90 (1.45)15.93 (2.34)9.91 (1.81)11.68 (1.58)14.34 (2.61)3.1 (1.48) Top-100012.21 (1.70)17.52 (2.29)15.04 (2.87)16.64 (4.07)16.81 (1.40)9.03 (2.29) Top-200012.57 (2.61)20.35 (1.40)14.51 (1.34)14.69 (1.61)18.76 (2.20)15.22 (4.35) Top-400013.27 (2.26)19.65 (0.74)16.46 (1.01)17.35 (4.18)18.76 (2.29)16.81 (1.25) Top-600013.63 (1.61)21.06 (2.53)16.28 (2.91)16.64 (2.61)18.94 (2.70)15.75 (2.58) Top-5 Real only18.58 (3.00)33.63 (3.70)24.60 (5.43)29.91 (2.68)26.19 (2.68)12.74 (2.84) Top-100027.26 (3.39)35.58 (3.39)32.04 (1.92)29.91 (1.45)33.81 (2.29)17.52 (4.26) Top-200030.09 (1.98)37.35 (2.37)33.63 (2.65)33.10 (3.99)36.81 (3.40)25.13 (2.70) Top-400030.09 (4.29)40.35 (4.08)33.98 (2.97)34.16 (2.84)35.40 (4.15)28.14 (0.97) Top-600033.45 (4.70)39.65 (2.29)32.74 (1.77)33.81 (1.58)36.46 (3.78)29.38 (2.45) Top-10 Real only28.50 (3.67)43.01 (2.63)34.87 (4.75)38.41 (1.34)35.93 (3.17)19.12 (5.18) Top-100038.41 (3.94)47.61 (3.30)44.25 (2.94)39.82 (1.77)46.19 (3.03)26.02 (2.55) Top-200041.59 (5.16)47.26 (2.13)43.89 (2.55)43.01 (1.34)46.02 (1.08)33.98 (2.25) Top-400040.53 (3.45)51.33 (3.00)43.72 (2.70)44.25 (3.00)46.19 (1.58)34.87 (1.61) Top-600042.65 (2.68)49.20 (1.48)43.89 (2.47)46.73 (1.15)48.67 (2.26)37.70 (2.03) Top-30 Real only54.34 (2.39)64.42 (1.92)58.23 (5.06)60.88 (2.02)58.41 (3.81)42.30 (4.40) Top-100062.30 (2.22)70.62 (3.03)64.07 (3.99)63.01 (2.83)70.62 (2.20)49.56 (2.65) Top-200062.12 (3.15)70.62 (3.39)64.60 (3.59)66.73 (3.99)69.56 (2.70)57.17 (4.18) Top-400063.36 (2.91)70.09 (1.48)67.96 (1.58)64.42 (2.37)70.80 (1.08)57.35 (2.02) Top-600063.19 (3.94)68.67 (3.23)63.01 (2.89)66.37 (2.58)68.50 (2.39)61.24 (3.05) Table S6. Top-k accuracies (%) across backbones and synthetic cutoffs of FastGAN samples. ACC (%)Top-nResNetDenseNetFaceNetVGGSwin-TCLIP Top-1 Real only6.90 (1.45)15.93 (2.34)9.91 (1.81)11.68 (1.58)14.34 (2.61)3.1 (1.48) Top-10008.50 (1.48)13.27 (3.13)6.55 (2.84)7.26 (2.37)10.44 (1.70)1.42 (1.34) Top-20006.55 (2.13)10.44 (0.40)5.13 (0.74)6.37 (3.67)8.50 (2.63)1.06 (1.15) Top-40004.07 (1.48)8.14 (1.45)4.96 (1.19)7.96 (2.26)9.73 (1.40)0.71 (0.74) Top-60004.78 (0.79)9.73 (2.08)4.78 (1.34)5.49 (2.61)9.20 (1.84)1.42 (1.01) Top-5 Real only18.58 (3.00)33.63 (3.70)24.60 (5.43)29.91 (2.68)26.19 (2.68)12.74 (2.84) Top-100020.71 (1.34)27.79 (1.34)18.05 (2.97)18.94 (3.40)25.66 (1.88)5.49 (2.37) Top-200019.12 (2.31)23.54 (2.22)18.05 (1.73)18.76 (2.89)25.49 (4.03)6.90 (2.76) Top-400017.17 (2.31)25.49 (2.89)16.81 (1.25)20.71 (2.04)26.19 (3.57)5.49 (1.70) Top-600018.58 (1.98)25.49 (2.02)15.58 (1.94)21.95 (0.97)23.89 (4.42)6.19 (1.08) Top-10 Real only28.50 (3.67)43.01 (2.63)34.87 (4.75)38.41 (1.34)35.93 (3.17)19.12 (5.18) Top-100029.20 (3.59)37.70 (3.29)30.27 (3.51)29.03 (5.92)36.64 (2.22)11.15 (1.48) Top-200028.67 (4.18)34.34 (0.74)28.85 (3.79)29.38 (2.20)35.58 (3.51)11.68 (2.20) Top-400029.03 (2.68)35.93 (3.04)28.32 (2.50)31.33 (4.23)36.64 (4.58)11.15 (2.77) Top-600031.15 (3.15)35.93 (3.10)24.07 (2.29)32.39 (2.97)33.45 (3.62)13.63 (0.79) Top-30 Real only54.34 (2.39)64.42 (1.92)58.23 (5.06)60.88 (2.02)58.41 (3.81)42.30 (4.40) Top-100051.68 (2.04)61.59 (5.82)57.52 (0.63)55.22 (5.07)60.35 (1.92)34.51 (2.73) Top-200053.10 (3.54)58.41 (2.08)55.04 (4.44)53.27 (3.33)61.77 (4.90)32.39 (2.13) Top-400056.46 (3.51)60.00 (3.15)55.58 (1.70)59.12 (4.12)57.88 (3.40)32.57 (2.46) Top-600055.75 (3.19)54.34 (3.68)55.40 (2.31)60.71 (1.94)56.46 (1.15)38.23 (5.78) Synthetic scaling effect The relationship between real-only and synthetic-augmented performance across Top-k settings and Top-n synthetic cutoffs is shown in Figure S8 and Figure S9. These plots visualize the downstream classification accuracies using DreamBooth and FastGAN generated samples, respectively, across six backbone models. Figure S8. Top-k accuracy comparison using DreamBooth-generated data. Each subplot shows Top-1, Top-5, Top-10, and Top-30 accuracy across synthetic cutoffs for six backbone models. DreamBooth augmentation improves performance across most settings. Figure S9. Top-k accuracy comparison using FastGAN-generated data. Each subplot shows Top-1, Top-5, Top-10, and Top-30 accuracy across synthetic cutoffs for six backbone models. Compared to DreamBooth, FastGAN augmentation results in less consistent or degraded performance across most settings. F.2. Few-shot learning Based on prior results, we focus our few-shot learning analysis on DreamBooth-augmented data. This choice is motivated by the consistent advantages of DreamBooth samples over FastGAN in terms of semantic fidelity, preservation of disease- relevant facial traits, and interpretability. These properties are especially important in ultra-low-shot settings, where maxi- mizing the realism and diagnostic relevance of synthetic data is critical for generalization. Table S7 summarizes performance across multiple backbones under synthetic data augmentation. Each pair of rows shares the same pretrained backbone, where the first row (e.g., ResNet) uses only real training data, and the second row (e.g., ResNet (dream)) incorporates DreamBooth- generated synthetic images for data augmentation. Table S7. Few-shot classification accuracies (%) under 5-way, 10-way, and 15-way settings with 1-shot and 5-shot support and query sets. ACC (%)5-way10-way15-way 1-shot5-shot1-shot5-shot1-shot5-shot ResNet24.18 (2.56)â18.15 (2.58)â17.99 (1.26)â ResNet (dream)25.72 (1.62)33.62 (1.91)22.21 (2.60)22.63 (2.11)18.22 (1.91)20.04 (2.77) DenseNet26.20 (2.01)â17.36 (1.66)â17.30 (3.45)â DenseNet (dream)29.88 (1.51)33.40 (2.02)19.79 (3.19)24.43 (3.82)16.06 (2.77)18.35 (2.24) FaceNet25.16 (4.89)â12.93 (2.33)â8.03 (1.40)â FaceNet (dream)23.60 (4.06)28.30 (2.92)13.17 (2.33)17.31 (3.47)9.07 (1.44)12.17 (1.24) VGG21.54 (3.72)â10.82 (2.14)â9.05 (2.20)â VGG (dream)21.02 (5.85)26.76 (3.42)13.67 (2.08)13.68 (1.94)9.59 (1.48)11.20 (1.59) Swin-T22.24 (2.88)â10.94 (2.69)â8.03 (2.41)â Swin-T (dream)26.72 (4.34)25.78 (4.91)13.30 (1.56)13.43 (1.94)9.57 (2.57)8.82 (0.97) CLIP23.48 (5.28)â12.33 (1.96)â7.30 (1.90)â CLIP (dream)22.30 (2.63)24.80 (3.48)12.14 (1.98)13.04 (2.08)8.66 (2.82)8.11 (1.32) F.3. Observations Standard supervised classification The results in Table S5 and Table S6 highlight a clear distinction in the effectiveness of different generative models for data augmentation in classifications. DreamBooth consistently enhances model performance across Top-k accuracies and Top-n settings, demonstrating its ability to produce high-quality, semantically aligned images that support generalization. In contrast, FastGAN shows less consistent gains and, in some cases, degrades performance as more synthetic data is addedâindicating a lower signal-to-noise ratio in the generated samples. These findings underscore the importance of not only using synthetic data, but selecting the right generation method. Few-shot learning Table S7 further illustrates how the utility of synthetic augmentation depends on the interaction between model architecture and data quality. While DreamBooth-generated samples consistently improve performance, gains vary across backbones, with CNNs such as DenseNet and ResNet benefiting more than transformer-based models like Swin Transformer or CLIP. Additionally, the performance gap between 1-shot and 5-shot conditions emphasizes the value of even modest increases in labeled support. These results suggest that few-shot learning with synthetic data is most effective when both model and augmentation strategy are jointly optimized for the task. Overall, high semantic fidelity, structural consistency, and disease-relevant visual features appear essential for synthetic augmentation to benefit rare disease classification, especially in ultra-low-shot regimes. G. VLM-based report generation G.1. Prompt design To ensure consistency and clinical interpretability of the generated phenotype reports, we designed a structured prompt tailored to the rare disease diagnosis task. The prompt guides the multimodal vision language model (VLM) to assume the role of a professional clinical geneticist and produce structured, concise diagnostic reports based on input facial images. The prompt emphasizes anatomical coverage, medical reasoning, and alignment with the supervised landmark annotations used in other parts of our study. A prompt template is shown in Figure S10. Figure S10. Prompt template for VLM-based report generation. G.2. Report evaluation To evaluate the semantic consistency of synthetic facial images, we used a multimodal large language model (VLM) to generate structured phenotype reports for real and DreamBooth-generated images. This analysis assesses whether synthetic faces reflect clinically meaningful traits of their target disease class. We excluded FastGAN images, as they are generated unconditionally and lack one-to-one correspondence with real patients. Since the VLM-based evaluation depends on matched image pairs for assessing regional semantic consistency, it is not applicable to class-agnostic generative models. To illustrate this evaluation process, we present selected examples spanning high and low semantic similarity scores. For each example, we compare the phenotype reports generated from a real image and a DreamBooth-generated synthetic image of the same disease class. Region-level descriptions are extracted from both reports. Matching phenotype terms are highlighted in green, while contradictory or inconsistent terms are shown in red. The examples cover a range of semantic similarity levels, allowing qualitative inspection of both faithful and divergent generations. Detailed comparisons are shown in Figure S11, S12, S13, and S14. Figure S11. High-Similarity Case. The real and synthetic reports both describe ptosis, a broad nose, and slighly open mouth appearance. Minor deviations in phrasing are present, but the overall diagnostic impression remains aligned. Figure S12. High-Similarity Case. Side-by-side comparison of phenotype reports from a real and synthetic image of the same disease class. The synthetic report closely mirrors the real one, especially in descriptions of palpebral fissures, ptosis, nasal shape, and lip structure. Figure S13. Mixed-Agreement Case. Although there is some alignment in features like nasal shape and bridge width, the synthetic report diverges in descriptors of lip configuration and diagnostic suggestion, showing partial inconsistency. Figure S14. Low-Similarity Case. The real report describes normal symmetry, while the synthetic report notes asymmetry and deviation, showing the VLM interprets synthetic structure differently. BioBERT semantic similarity analysis We report the detailed results of the semantic similarity analysis between real and synthetic phenotype descriptions. Similarity scores were computed using BioBERT embeddings across five facial regions and an overall report segment. Table S8 summarizes the mean and standard deviation of cosine similarities across the dataset. The highest alignment was observed in the nose and eye regions, while the mouth/lips showed slightly greater variability. TF-IDF comparison To complement the analysis presented above, we conducted a parallel evaluation using a traditional text similarity method based on TF-IDF [33] (Table S8). For each facial region and the overall report, we computed cosine similarity between real and synthetic descriptions using scikit-learn [37]âs TfidfVectorizer. Table S8. Semantic similarity and TF-IDF-based semantic similarity across five facial regions. RegionBioBERT Similarity ScoreTF-IDF Similarity Score Overall0.8404 (0.0748)0.7630 (0.0707) Left Eye0.7485 (0.1228)0.4364 (0.1358) Right Eye0.7535 (0.1423)0.4578 (0.1352) Nose0.7712 (0.1344)0.4315 (0.1432) Mouth/Lips0.7355 (0.1376)0.4612 (0.1321) These findings support the need for domain-specific semantic models when evaluating medical text. TF-IDFâbased com- parisons fail to fully capture conceptual similarity in phenotype language. While overall similarity trends are consistent, the lower region-wise scores highlight the limitations of surface-level lexical methods when applied to clinical narratives. G.3. Uncertainty and robustness analysis Figure S15 complements the uncertainty analysis in the main text by visualizing the distribution of uncertainty scores across all samples. 75% of the cases exhibit low uncertainty (< 0.14), indicating stable and consistent phenotype descriptions under stochastic sampling. Figure S15. Distribution of uncertainty scores across all cases. To further assess robustness, Table S9 reports cross-model phenotype similarity between Qwen2.5-VL and LLaVA-NeXT. The results show consistent similarity ranges across real and synthetic images, supporting that the evaluation is not sensitive to the choice of VLM. Table S9. Cross-model phenotype similarity. LLaVA (Real)LLaVA (Synthetic) Qwen (Real)0.7053 (0.0711)0.7176 (0.0691) Qwen (Synthetic)0.7204 (0.0726)0.7355 (0.0730) H. Regional analysis and potential bias To assess potential bias in both classification performance and synthetic data evaluation, we analyze model behavior across geographic regions. However, population-level demographic attributes (e.g., ethnicity, ancestry, or skin tone) are not available in our dataset, as web-scraped rare disease case reports rarely provide standardized annotations. As a proxy, we stratify the data by geographic region (the only consistently recoverable attribute) and group samples into four regions: Africa, Americas, Asia, and Europe. Regional diagnosis performance. We report Top-k accuracy for supervised learning across regions in Figure S16a. While minor variations are observed at lower k, performance trends are broadly consistent across regions and converge as k in- creases, indicating similar generalization behavior as the results in main text. (a) Regional diagnosis results across Top-k accuracies.(b) Regional phenotype similarity comparison. Figure S16. Regional analysis across geographic groups. Regional phenotype similarity. We further evaluate phenotype similarity across regions using VLM-generated reports (Figure S16b). Realâsynthetic similarity is comparable to realâreal similarity across all regions within statistical uncertainty, suggesting that synthetic data preserves phenotype characteristics without introducing substantial regional bias. Notably, the Americas region shows slightly higher similarity scores, which may reflect a larger representation of cases from this region in the dataset. However, the overall consistency across regions supports the generalizability of our findings. Limitations. We note that geographic region is only a coarse proxy and does not directly correspond to demographic attributes such as skin tone or ancestry. In addition, potential collection bias in publicly available case imagery may influence both classification and generative models. Future work should prioritize the collection of more diverse and well-annotated datasets to enable more granular analysis of demographic bias and ensure equitable performance across all patient populations and the development of synthetic data generation methods that explicitly account for demographic diversity. I. Statistical reporting details All reported results in this paper include 1-sigma error bars, expressed as the standard deviation in parentheses (e.g., 6.90 (1.45)), computed over 5-fold cross-validation. The primary source of variability is the random train/test split across folds. For each configuration, we evaluate performance on the held-out fold and report the sample mean and standard deviation across the five runs. We assume approximately normally distributed metrics across folds, which is common in classification evaluation under low-data regimes. No additional sources of randomness (e.g., weight initialization or stochastic sampling) are varied unless explicitly stated. Standard deviation is calculated using the unbiased estimator (i.e., Besselâs correction): std(x 1 ,...,x n ) = v u u t 1 nâ 1 n X i=1 (x i â Ìx) 2 This computation is implemented via NumPy [16]âs np.std(..., ddof=1) function. J. Hyperparameter settings and training details We summarize the key hyperparameters used in our benchmark experiments across three experimental settings in Ta- ble S10, S11, S12, and S13. J.1. Standard supervised classification Table S10. Hyperparameters for standard supervised classification experiments. ParameterValue OptimizerAdam Learning rate1e â4 Batch size32 Training folds5-fold class split Number of epochs50 Loss functionCrossEntropyLoss Weight decay1e â4 Image size224Ă 224 Pretrained backboneImageNet for ResNet, DenseNet, Swin-T, CLIP VggFace2 for FaceNet VggFace for VGG J.2. Few-Shot learning Table S11. Hyperparameters for Prototypical Networks. ParameterValue OptimizerAdam Learning rate1e â3 Batch size1 Training folds5-fold class split Episodes per fold600 train / 100 val / 200 test Loss functionCrossEntropyLoss Distance metricEuclidean Distance J.3. Synthetic data generation Table S12. Training settings for DreamBooth per disease class. ParameterValue / Description Base modelSG161222/RealisticVisionV5.1noVAE Training settingClass-conditioned (per disease) Text prompt "a child with [DISEASE] disease" Training steps800 per class Batch size1 Learning rate1e â6 Image resolution512Ă 512 Mixed precisionfp16 Safety checkerNSFW filter Table S13. Training settings for FastGAN on pooled real images. ParameterValue / Description ArchitectureOriginal FastGAN (skip-layer excitation) Training settingUnconditional (no class labels) Iterations80,000 Batch size8 OptimizerAdam Learning rate2e â4 Image resolution512Ă 512 J.4. Hardware and compute resources Standard supervised classification, few-shot learning, DreamBooth, and FastGAN models were all implemented in PyTorch [36] and trained using a single NVIDIA A100 GPU with 80GB of memory (also used for VLM-based report generation). Specifically, standard supervised classification and few-shot learning required approximately 240 and 550 minutes of training time, respectively, for downstream analysis. DreamBooth required 30â50 minutes of training per class, depending on the number of input images and prompt complexity, while FastGAN training took approximately 15 hours to converge.