Paper deep dive
SkinGPT-X: A Self-Evolving Collaborative Multi-Agent System for Transparent and Trustworthy Dermatological Diagnosis
Zhangtianyi Chen, Yuhao Shen, Florensia Widjaja, Yan Xu, Liyuan Sun, Zijian Wang, Hongyi Chen, Wufei Dai, Juexiao Zhou
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/31/2026, 1:38:41 AM
Summary
SkinGPT-X is a multimodal collaborative multi-agent system designed for dermatological diagnosis. It features a self-evolving memory mechanism (EvoDerma-Mem) that allows the system to continuously accumulate historical case data and refine diagnostic guidelines without parameter retraining. The system outperforms state-of-the-art LLMs on multiple benchmarks, including rare skin disease datasets, by simulating a human-like diagnostic workflow involving visual findings extraction, hypothesis generation, and evidence-based cross-verification.
Entities (5)
Relation Signals (3)
SkinGPT-X → utilizes → EvoDerma-Mem
confidence 100% · SkinGPT-X... integrated with a self-evolving dermatological memory mechanism.
SkinGPT-X → outperforms → PanDerm
confidence 95% · SkinGPT-X consistently outperforms baseline models across all 4 evaluation metrics and datasets.
EvoDerma-Mem → stores → Historical Case Graph Database
confidence 90% · The Summarize Agent ensures self-evolving agent memory by integrating new confirmed cases into the Dynamic Repository
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While recent advancements in Large Language Models have significantly advanced dermatological diagnosis, monolithic LLMs frequently struggle with fine-grained, large-scale multi-class diagnostic tasks and rare skin disease diagnosis owing to training data sparsity, while also lacking the interpretability and traceability essential for clinical reasoning. Although multi-agent systems can offer more transparent and explainable diagnostics, existing frameworks are primarily concentrated on Visual Question Answering and conversational tasks, and their heavy reliance on static knowledge bases restricts adaptability in complex real-world clinical settings. Here, we present SkinGPT-X, a multimodal collaborative multi-agent system for dermatological diagnosis integrated with a self-evolving dermatological memory mechanism. By simulating the diagnostic workflow of dermatologists and enabling continuous memory evolution, SkinGPT-X delivers transparent and trustworthy diagnostics for the management of complex and rare dermatological cases. To validate the robustness of SkinGPT-X, we design a three-tier comparative experiment. First, we benchmark SkinGPT-X against four state-of-the-art LLMs across four public datasets, demonstrating its state-of-the-art performance with a +9.6% accuracy improvement on DDI31 and +13% weighted F1 gain on Dermnet over the state-of-the-art model. Second, we construct a large-scale multi-class dataset covering 498 distinct dermatological categories to evaluate its fine-grained classification capabilities. Finally, we curate the rare skin disease dataset, the first benchmark to address the scarcity of clinical rare skin diseases which contains 564 clinical samples with eight rare dermatological diseases. On this dataset, SkinGPT-X achieves a +9.8% accuracy improvement, a +7.1% weighted F1 improvement, a +10% Cohen's Kappa improvement.
Tags
Links
- Source: https://arxiv.org/abs/2603.26122v1
- Canonical: https://arxiv.org/abs/2603.26122v1
Trouble viewing inline? Open PDF directly →
Full Text
91,889 characters extracted from source content.
Expand or collapse full text
1 SkinGPT-X: A Self-Evolving Collaborative Multi-Agent System for Transparent and Trustworthy Dermatological Diagnosis Zhangtianyi Chen 1,† , Yuhao Shen 1,† , Florensia Widjaja 1,† , Yan Xu 2 , Liyuan Sun 3 , Zijian Wang 1 , Hongyi Chen 1 , Wufei Dai 1 , Juexiao Zhou 1,∗ Abstract—While recent advancements in Large Language Models have significantly advanced dermatological diagnosis, monolithic LLMs frequently struggle with fine-grained, large-scale multi-class diagnostic tasks and rare skin disease diagnosis owing to training data sparsity, while also lacking the interpretability and traceability essential for clinical reasoning. Although multi-agent systems can offer more transparent and explainable diagnostics, existing frameworks are primarily concentrated on Visual Question Answering and conversational tasks, and their heavy reliance on static knowledge bases restricts adaptability in complex real-world clinical settings. Here, we present SkinGPT-X, a multimodal collaborative multi-agent system for dermatological diagnosis integrated with a self-evolving dermatological memory mechanism. By simulating the diagnostic workflow of dermatologists and enabling continuous memory evolution, SkinGPT-X delivers transparent and trustworthy diagnostics for the management of complex and rare dermatological cases. To validate the robustness of SkinGPT-X, we design a three-tier comparative experiment. First, we benchmark SkinGPT-X against four state-of-the-art LLMs across four public datasets, demonstrating its state-of-the-art performance with a +9.6% accuracy improvement on DDI31 and +13% weighted F1 gain on Dermnet over the state-of-the-art model. Second, we construct a large-scale multi-class dataset covering 498 distinct dermatological categories to evaluate its fine-grained classification capabilities. Finally, we curate the rare skin disease dataset, the first benchmark to address the scarcity of clinical rare skin diseases which contains 564 clinical samples with eight rare dermatological diseases. On this dataset, SkinGPT-X achieves a +9.8% accuracy improvement, a +7.1% weighted F1 improvement, a +10% Cohen’s Kappa improvement. The self-evolving agent memory enables the continuous accumulation of historical cases and the iterative evolution of diagnostic guidelines, which significantly enhances the system’s reasoning depth across an expanding range of skin diseases. In summary, SkinGPT-X represents the first multimodal collaborative multi-agent-driven dermatological diagnosis system empowered by a self-evolving memory mechanism, enabling transparent and trustworthy diagnosis, particularly for rare skin diseases affected by insufficient training data. Index Terms—Dermatology, Multi-agent system, Large language model ✦ 1 INTRODUCTION Skin diseases exert a substantial global health burden, pro- foundly impacting not only physical well-being but also the psychosocial quality of life for millions [1], [2]. Recent updates from the Global Burden of Disease study further underscore that skin and subcutaneous diseases remain leading causes of non-fatal disability worldwide [3]. Despite this growing recognition of skin health as a pillar of over- all well-being, the accessibility of specialized care remains severely constrained by a systemic shortage of licensed prac- 1 School of Data Science, The Chinese University of Hong Kong, Shenzhen 2 Department of Dermatology, Tianjin Institute of Integrative Dermatology, Tianjin Academy of Traditional Chinese Medicine Affiliated Hospital, Tianjin 300120, China 3 Department of Dermatology, Beijing AnZhen Hospital, Capital Medical University, Beijing 100029, China.. † These authors contributed equally. ∗ Corresponding author. e-mail: juexiao.zhou@gmail.com titioners [4]. Consequently, the diagnostic benefits for most minor skin ailments are disproportionately low compared to the time and effort patients expend to visit these major institutions [5], [6]. As the result, teledermatology becomes more and more popular in order to expand the range of services available to medical professionals [7], [8]. Yet, while teledermatology addresses the issue of geographical distance, it does little to alleviate the absolute scarcity of expert labor [9]. In fact, the surge in digital consultations often overwhelms the existing clinical workforce [10]. This bottleneck has necessitated the integration of automated diagnostic tools [11] to facilitate triage and preliminary screening in dermatology. Artificial Intelligence (AI) is widely regarded as having immense potential to augment human expertise in specific medical domains [12]. Its application in dermatology holds profound significance in reducing healthcare costs and en- hancing diagnostic efficiency [13], [14], [15]. The break- arXiv:2603.26122v1 [cs.CV] 27 Mar 2026 2 through in Deep Learning (DL) has witnessed a paradigm shift, moving from traditional visual inspections to auto- mated, high-precision diagnostic systems [16]. Over the past decades, diverse DL architectures have been deployed across various dermatological domains. These include stan- dard Convolutional Neural Networks (CNNs) such as VGG- 16/19 [17], ResNet-50 [18], and Inception-V3 [19] for robust feature extraction, as well as lightweight models like Mo- bileNet V2 [20] and Xception [21] designed for mobile-based point-of-care diagnostics. These DL-based methods have been proposed for various dermatological domains like skin disease classification [22], [23], [24], [25], [26], [27], [28], skin cancer diagnosis [29], [30], [31], [32], melanoma detection [33], [34], [35], [36], [37], and analysis for different types of psoriasis [38]. Beyond classification, advanced frameworks involving Generative Adversarial Networks (GANs) [39] for data augmentation and hybrid deep learning approaches for skin lesion segmentation [40] have significantly improved the detection of malignant conditions. Despite achieving expert-level accuracy in controlled datasets, recent studies highlight critical ongoing challenges such as the need for cross-population generalization [41], handling class imbal- ance in rare skin diseases [42], [43], and the imperative for explainable AI (XAI) to foster clinical trust [32], [35]. Recently, Large Language Models (LLMs) have emerged as a transformative force in artificial intelligence. Compared to traditional DL models, LLMs exhibit superior multi- modal integration capabilities [44], [45], [46], [47], which can capture long-range dependencies between disparate data modalities, including clinical text, dermoscopic images, and omics data. The application of LLMs in dermatology extends beyond conventional skin cancer diagnosis [48], [49], [50], [51], [52] and clinical condition classification [53], [54], [55]. These models also demonstrate significant potential to augment traditional healthcare workflows by providing medication guidance [56] and addressing com- mon patient inquiries with high accessibility [57], [58], [59]. Furthermore, LLMs offer a promising solution for bridging the communication gap between clinicians and patients by translating intricate medical terminology into easily com- prehensible language [60], [61], [62]. This enables a more holistic understanding of complex pathologies [63], [64] and facilitates precision medicine by tailoring treatment schedules [65]. Furthermore, LLMs serve as pivotal tools for patient education and physician-patient interaction by simplifying complex medical reports into patient-friendly language [66], [67], [68]. Although LLMs have made sig- nificant progress, there are still two main limitations when applying them to clinical diagnosis. 1) Bottlenecks in High- Cardinality and Rare Disease Tasks: The static nature of fine-tuned knowledge bases proves inadequate for large- scale diagnostic tasks involving an extensive label space and the identification of rare diseases. These models often fail to generalize across such a vast and fine-grained dis- tribution of pathologies, falling short of the performance required for real-world clinical deployment [69]. 2) Insuffi- cient Traceability of Standalone Models: While standalone LLMs can provide linguistic explanations, their descriptive granularity often falls short of rigorous clinical standards. Such models lack the necessary diagnostic traceability and evidence-based reasoning required for high-stakes medical accountability [70]. Multi-Agent System represents a promising paradigm for overcoming these limitations by establishing a traceable chain of evidence through collaborative intelligence [71], [72]. By distributing complex reasoning across a network of specialized agents where each assumes distinct roles such as clinical analysis, evidence retrieval, or cross-verification, multi-agent system simulates the consultative process of a multi-disciplinary medical team [73], [74]. This synergetic workflow allows the system to decompose intricate diagnos- tic tasks into discrete, cross-validated stages [75], effectively transforming monolithic LLM outputs into structured, ver- ifiable, and evidence-based diagnostic pipelines [71]. Cur- rently, the use of multi-agent systems in healthcare is still in its infancy, largely limited to text-based conversations [74]. However, in real-world settings like teledermatology, the ability to process multi-modal data is of paramount impor- tance for diagnostic accuracy [76]. Furthermore, Retrieval- Augmented Generation (RAG) is extensively employed to equip these agents with specialized medical knowledge [77], [78]. Despite these efforts, static knowledge bases typically encapsulate only standardized standards from textbooks, which often prove inadequate when confronted with atyp- ical clinical manifestations or rare variant presentations in complex real-world scenarios [43]. To address these challenges, we propose a novel multi- agent system architecture named SkinGPT-X, inspired by authentic clinical diagnostic workflows. Our system harmo- nizes textbook-based expertise, empirical experience, and visual interpretation into a collaborative framework (Fig. 1b). First, we decouple the pre-diagnostic inference from visual feature findings extraction by employing two spe- cialized agents: one dedicated to extracting visual findings and the other to generating initial diagnostic hypothesis. To ensure foundational clinical accuracy, we employ a RAG module to anchor the initial reasoning process in evidence-based dermatological common sense [77]. Sec- ond, we propose a Self-Evolving Dermatological Diagnostic Memory(EvoDerma-Mem) mechanism, which empowers agents to autonomously refine their internal knowledge bases without the need for parameter retraining. Unlike 3 Disease Image Confirmed Diagnosis Historical Case Graph Database (Dynamic Repository) Vision Agent (Diagnosis-Guided Extraction) Summarize Agent (Category Experience Update) Combines New & Existing Similar Data New Cases Added Retrieve Similar Existing Cases Update Disease Category Summary Stored Cases Summarized Category BCC/AK Eczema Psoriasis KBC Btra Otherwise PRE-DIAGNOSIS HYPOTHESIS VISUAL FINDINGS solitary nodule central pigmentation ulceration ... DISEASE DIAGNOSIS AGENT VISION AGENT RAG (Retrieval Augmentation Generation) MEDICAL KNOWLEDGE BCC Guidelines Eczema Guidelines SUMMAR Lesion Characteristics Atypical Presentations Image Encoder Embedding Vector Historical Cases Database Key Findings: erythematous-pink macule Prior Diagnosis: BCC/AK Key Findings: erythematous plaque, crusted center Prior Diagnosis: Cellulitis Key Findings: erythematous plaque, crusted center Prior Diagnosis: SCC Key Findings: erythematous plaque, crusted center Prior Diagnosis: BCC/AK Key Findings: diffuse eruption, pink macules Prior Diagnosis: SCC 1 3 5 2 4 CASE-REVIEW AGENT Key Findings: The patient presents with a solitary, well- circumscribed, nodular, ulcerated, and crusted lesion on the anterolateral thigh. It has central dark brown to black pigmentation, a surrounding erythematous halo, and irregular, notched borders. The surface is necrotic, with ulceration and crusting at the periphery. There is no scaling, lichenification, or telangiectasia. The lesion is asymmetric, lacks internal structural patterns, and is consistent with a malignant or highly inflamed process. PrimaryDiagnosis: BCC/AK and other Malignant Lesions Evidence: The visual findings show a well-circumscribed, nodular, ulcerated, and crusted lesion with central pigmentation and erythematous halo, which aligns with the clinical features of Actinic Keratosis or Basal Cell Carcinoma. The presence of irregular borders, asymmetry, and lack of structural patterns (e.g., pigment network, globules) are concerning for malignancy. The model's high probability (87.80%) for Actinic Keratosis / Basal Cell Carcinoma and other Malignant Lesions is supported by the lesion's morphology and location (sun-exposed skin). Additionally, the patient's presentation matches several past cases of malignant lesions, particularly those with ulceration, necrosis, and erythematous margins. While some physically similar cases were classified as viral infections or pigmentation disorders, the clinical and histopathological features of the lesion are more consistent with a malignant process. The model's confidence is reasonable given the visual cues and the lack of features pointing toward benign conditions. Diagnosis Result Treatment Plan Online Verification (latest treatments) Standard Medical Literature & Textbooks Personal Clinical Experience & Knowledge Reserve Recall of Similar Historical Cases 0.88 0.38 0.28 0.16 0.13 0.10 Disease Image Confirmed Diagnosis Historical Case Graph Database (Dynamic Repository) Vision Agent (Diagnosis-Guided Extraction) Summarize Agent (Category Experience Update) Combines New & Existing Similar Data New Cases Added Retrieve Similar Existing Cases Update Disease Category Summary Stored Cases Summarized Category Diagnosis Result Treatment Plan Online Verification (latest treatments) Standard Medical Literature & Textbooks Personal Clinical Experience & Knowledge Reserve Recall of Similar Historical Cases Embedding Vector Self-Evolving Agent Memory Disease Image Confirmed Diagnosis Vision Agent (Diagnosis-Guided Extraction) Ground Truth Visual Features (Embeddings) Nature Language Description Summarize Agent (Category Experience Update) Combines New & Existing Similar Data Historical Case Graph Database (Dynamic Repository) New Cases Similar Existing Cases Update Diagnostic Guidelines Image Encoder Top-5 Similar Cases Image Question e.g. “What disease have I got and how can I treat it?” VALIDATED REPORT a b c a b Case-Review Agent BCC/AK Eczema Psoriasis Cellulitis Dermatitis Otherwise PRE-DIAGNOSIS HYPOTHESIS VISUAL FINDINGS solitary nodule central pigmentation ulceration ... Disease Diagnosis Agent Vision Agent RAG (Retrieval Augmentation Generation) MEDICAL KNOWLEDGE BCC Guidelines Eczema Guidelines DIAGNOSIS GUIDELINES Lesion Characteristics Atypical Presentations Key Findings: erythematous-pink macule Prior Diagnosis: BCC/AK Key Findings: erythematous plaque, crusted center Prior Diagnosis: Cellulitis Key Findings: erythematous plaque, crusted center Prior Diagnosis: SCC Key Findings: erythematous plaque, crusted center Prior Diagnosis: BCC/AK Key Findings: diffuse eruption, pink macules Prior Diagnosis: SCC 1 3 5 2 4 Key Findings: The patient presents with a solitary, well- circumscribed, nodular, ulcerated, and crusted lesion on the anterolateral thigh. It has central dark brown to black pigmentation, a surrounding erythematous halo, and irregular, notched borders. The surface is necrotic, with ulceration and crusting at the periphery. There is no scaling, lichenification, or telangiectasia. The lesion is asymmetric, lacks internal structural patterns, and is consistent with a malignant or highly inflamed process. PrimaryDiagnosis: BCC/AK and other Malignant Lesions Evidence: The visual findings show a well-circumscribed, nodular, ulcerated, and crusted lesion with central pigmentation and erythematous halo, which aligns with the clinical features of Actinic Keratosis or Basal Cell Carcinoma. The presence of irregular borders, asymmetry, and lack of structural patterns (e.g., pigment network, globules) are concerning for malignancy. The model's high probability (87.80%) for Actinic Keratosis / Basal Cell Carcinoma and other Malignant Lesions is supported by the lesion's morphology and location (sun-exposed skin). Additionally, the patient's presentation matches several past cases of malignant lesions, particularly those with ulceration, necrosis, and erythematous margins. While some physically similar cases were classified as viral infections or pigmentation disorders, the clinical and histopathological features of the lesion are more consistent with a malignant process. The model's confidence is reasonable given the visual cues and the lack of features pointing toward benign conditions. 0.88 0.38 0.28 0.16 0.13 0.10 Image Question e.g. “What disease have I got and how can I treat it?” Stored Cases Validated Report Human Cognitive SynthesisHuman Cognitive SynthesisHuman Cognitive Synthesis New Cases Added Summarized Category Fig. 1. a, The process of clinicians formulating a comprehensive Diagnosis Report: Clinicians formulate a Diagnosis Report by integrating three core pillars of medical intelligence: Personal Clinical Experience, Standard Medical Literature, and the Recall of Similar Historical Cases. This holistic synthesis, further augmented by Online Verification of the latest treatments, enables a validated Treatment Plan. b, The Architecture of the SkinGPT-X System: Upon receiving a disease image and a query, the Vision Agent extracts fine-grained Visual Findings, while the Diagnosis Diagnosis Agent performs an initial Pre-diagnosis Hypotheses. This process is anchored by the RAG, which retrieves Local Medical Knowledge to ensure evidence-based reasoning. To incorporate empirical experience, the system utilizes the self-evolving agent memory, where the Top-5 Similar Cases and Diagnostic Guidelines are retrieved from the Historical Case Graph Database. The visual features, retrieved guidelines, and historical precedents are synthesized by the Case-Review Agent, which conducts a rigorous cross-reference to produce a validated Report. The Summarize Agent ensures self-evolving agent memory by integrating new confirmed cases into the Dynamic Repository, iteratively distill the diagnostic guidelines without retraining. 4 traditional systems that rely on static textbooks, SkinGPT-X iteratively synthesizes diagnostic experiences from encoun- tered cases. As the volume and diversity of cases grow, the system generates comprehensive diagnostic guidelines that encapsulate both prototypical disease markers and rare clinical variants. This evolutionary approach ensures that the system continuously broadens its understanding of dermatological conditions, offering superior scalability and coverage compared to static knowledge base. During the in- ference phase, the system retrieves visually similar historical cases and corresponding guidelines from the agent memory. SkinGPT-X then cross-references these insights with the current visual features to conduct a rigorous review of the preliminary diagnosis. This architecture ensures that fine- grained disease subcategories and rare pathologies receive equitable attention. Furthermore, the self-generated agent memory acts as a corrective layer, enabling the system to de- tect subtle diagnostic oversights and evaluate the possibility of rare conditions, ultimately producing a more traceable, transparent, and trustworthy diagnostic report. 2 RESULTS 2.1 The overall design of SkinGPT-X SkinGPT-X operates as a self-evolving multi-agent system that bridges raw visual data with traceable clinical evidence as show in Fig. 1. In the Diagnostic Phase, multiple agents collaborate to analyze patient images from diverse per- spectives, including visual findings, preliminary diagnostic hypotheses and retrieval book-based medical knowledge. These insights are then audited by a Case-Review Agent, which is powered by a self-evolving agent memory to produce a final, validated report. The system’s Evolutionary Phase ensures continuous learning: new cases are archived as memory nodes, and upon reaching a knowledge density threshold, a Summarization Agent automatically synthe- sizes these historical cases into high-level clinical guidelines, emulating the expertise accumulation of a human dermatol- ogist. To rigorously evaluate the diagnostic efficacy, robust- ness, and interpretability of SkinGPT-X, we designed a three-tiered experimental framework ranging from common pathologies to complex rare diseases. In Section 2.2, we conducted comprehensive benchmarking against three gen- eral medical LLMs and one fine-tuned foundation model (PanDerm [79]) across 4 public datasets (Dermnet [80], HAM10000 [81], DDI [82], and Fitzpatrick-17k [83]). Our findings indicate that SkinGPT-X not only excels in standard metrics like Accuracy (ACC) and Weighted F1 but also significantly outperforms baselines in Matthews Correlation Coefficient (MCC) and Cohen’s Kappa, suggesting its ex- ceptional potential in addressing the fine-grained diagnosis task. Motivated by this observation, we sought to further explore the system’s performance under high-cardinality label spaces. In Section 2.3, we reconstructed the Dermnet dataset into Dermnet498, shifting from 23 broad super- classes to 498 granular sub-classes based on hierarchical metadata. Based on it, we systematically explored how SkinGPT-X maintains its stability as categorical complex- ity intensifies. We also curated a novel Rare Skin Disease Dataset (RSDD) based on the Rare Disease Diagnosis and Treatment Guidelines published by the National Health Commission of China in Section 2.4. This benchmark served to stress-test the model’s few-shot reasoning capabilities. Finally, to validate the core architectural innovation of our system, we conducted ablation studies and quantitative evaluations focusing on the EvoDerma-Mem in Section 2.5. These experiments confirms the efficacy of the the proposed framework and the EvoDerma-Mem mechanism in synthe- sizing diagnostic rationales and evolving through historical case accumulation. 2.2 SkinGPT-X Achieves State-of-the-Art Performance Against Four Baselines Across Four Datasets To evaluate the performance of SkinGPT-X, we conducted comprehensive comparative experiments across 4 bench- mark datasets, including Dermnet, HAM10000, DDI31, and Fitzpatrick-17k. To provide a comprehensive view of their diagnostic efficacy, we assessed the models using 4 metrics: ACC, Weighted F1, MCC, and Cohen’s Kappa. Among these models, PanDerm was fine-tuned on each specific dataset for multi-class classification. In contrast, MedGemma [84], Qwen3-VL [85] and Hulu-Med [86], as general-purpose LLMs, conducted closed-ended dermatological assessments by selecting the most probable category from the label space of each dataset. As illustrated in Fig. 2a, SkinGPT-X consis- tently outperforms baseline models across all 4 evaluation metrics and datasets. In the Dermnet dataset (n = 19, 559), SkinGPT-X achieves a Weighted F1 of 87.9% and an MCC of 79.9%, significantly surpassing existing methods with high statistical confidence (P < 0.001). In the HAM10000 dataset (n = 10, 015), SkinGPT-X achieves the best performance in ACC (86.6%), Weighted F1 (86.4%), MCC (77.4%), and Kappa (77.2%). Notably, SkinGPT-X shows a substantial lead in the DDI31 dataset which typically presents a greater challenge due to diverse skin tones, reaching a Weighted F1 of 29.2%, an ACC of 35.8%, an MCC of 30.4%, and a Kappa of 25.5%, markedly higher than the leading baselines. Furthermore, in the Fitzpatrick-17k dataset (n = 16, 577), SkinGPT-X maintains its competitive edge, achieving the highest scores across all 4 metrics. These results collectively 5 MedGemmaSkinGPT-X The image shows a case of severe psoriasis, characterized by erythematous, scaly plaques with visible pustules, particularly on the hands. This indicates a significant inflammatory skin condition. The standard treatment for severe psoriasis includes a combination of topical corticosteroids, vitamin D analogs (calcipotriene), and phototherapy (UVB or PUVA). .... Adjunctive therapies include emollients for skin hydration and topical calcineurin inhibitors for mild to moderate psoriasis. Primary Diagnosis: Psoriasis Visual Findings: The image shows multiple, discrete, raised, irregularly shaped nodules and papules on the dorsal hand and fingers. Lesions are erythematous to violaceous with hyperpigmentation, central ulceration/crusting, irregular/poorly defined borders, and heterogeneous surface texture (scaling, crusting, ulceration, induration). Marked asymmetry in distribution. No prominent telangiectasia or lichenification. Case Reviewing: The visual findings of multiple, irregularly shaped, erythematous to violaceous nodules and papules with scaling, crusting, ulceration, and poorly defined borders are inconsistent with the standard description of erythema elevatum diutinum (smooth, glistening, well-circumscribed, no scaling/crusting/ulceration). While the model predicts erythema elevatum diutinum (23.69%), the description aligns better with the potential presentations of porphyria (15.44%), which can include papules, nodules, plaques with scaling, crusting, and irregular borders. Past cases of porphyria show varied morphologies, including scaly plaques and erosions, supporting the possibility of an atypical presentation. Factitial dermatitis (2.80%) is also a consideration given the ulceration and irregular morphology, but the model ranks it lower. Lupus erythematosus (4.87%) is less likely given the lack of typical features like atrophy or prominent telangiectasia. Milia (3.43%) is clearly excluded by the morphology (nodules, papules, plaques, ulceration vs. small, smooth papules). Primary Diagnosis: Porphyria Hulu-Med The image shows a case of psoriasis vulgaris, plaque type, characterized by erythematous, scaly plaques with visible thickening, particularly on areas prone to plaque formation. This indicates a chronic inflammatory skin condition with well-demarcated lesions. The Standard treatment for psoriasis vulgaris includes a combination of topical corticosteroids, vitamin D analogs .....Adjunctive therapies include moisturizers for skin hydration and sun protection to prevent exacerbation. Primary Diagnosis: Psoriasis Vulgaris 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.1 0.2 0.3 0.4 0.5 0.0 0.2 0.4 0.6 0.8 1.0 a b Dermnet (c=23; n=19,559)HAM10000 (c=7; n=10,015) Fitzpatrick17k (c=114; n=16,577 )DDI31 (c=31; n=647) MedGemmaHulu-MedQwen3-VLPanDermSkinGPT-X ACCWeighted F1MCCKappaACCWeighted F1MCCKappa ACCWeighted F1MCCKappaACCWeighted F1MCCKappa P< 0.001P< 0.001P< 0.001P< 0.001 0.750 0.790 0.879 0.749 0.728 0.799 P< 0.001P< 0.001P< 0.001P< 0.05 P< 0.001P< 0.001P< 0.001P< 0.001 0.735 0.780 0.847 0.8660.864 0.850 0.752 0.774 0.749 0.772 0.262 0.358 0.292 0.234 0.211 0.303 0.207 0.255 0.671 0.707 0.725 0.668 0.667 0.703 0.667 0.703 P< 0.001P< 0.001P< 0.001P< 0.001 Ground Truth: Prophyria Fig. 2. a, Performance comparison of SkinGPT-X versus four state-of-the-art models (MedGemma, Hulu-Med, Qwen3-VL, and PanDerm) across four benchmark skin disease datasets. c indicates the class number; n represents the data size. Metrics include ACC, Weighted F1, MCC, and Cohen’s Kappa. The error bars represent 95% CIs, and P values were calculated using a two-sided t-test to indicate statistical significance between SkinGPT-X and the second-best model. b, Case study of diagnostic case reviewing and visual findings generation. The panels display the text- based outputs from MedGemma, Hulu-Med, and SkinGPT-X for representative clinical cases. The results illustrate that SkinGPT-X provides more reliable and transparent diagnostics by reviewing current clinical findings with its self-evolving agent memory. underscore SkinGPT-X’s generalization and diagnostic ac- curacy across varying scales and complexities of dermato- logical data. 2.3 SkinGPT-X is Effective in High-Dimensional Classi- fication Scenarios As illustrated in Fig. 3a, while the original Dermnet dataset partitions samples into 23 categories, we found that these broad categories often grouped together clinically distinct diseases, leading to unclear definitions of specific condi- tions. Specifically, labels such as Psoriasis pictures Lichen Planus and related diseases and Tinea Ringworm Candidiasis and other Fungal Infections aggregate diverse conditions. To address, we introduced Dermnet498, a benchmark that re-categorizes the original images into 498 fine-grained classes based on their intrinsic, sample-specific metadata. This refinement process deconstructs broad categories into clinically sub-classes. For instance, the category of Psoriasis 6 23225272353498 0.5 0.6 0.7 0.8 0.9 23225272353498 0.5 0.6 0.7 0.8 0.9 23225272353498 0.5 0.6 0.7 0.8 0.9 Basal Cell Carcinoma (535) Lichen Planus (246) Molluscum Contagiosum (242) Alopecia Areata (208) Warts Common (162) Other 494 Categories (16,924) 0.0 0.2 0.4 0.6 0.8 Performance Comparison on Dermnet498 ACCMCCWeighted F1Kappa MedGemma Hulu-Med Qwen3-VL SkinGPT-X PanDerm P< 0.001P< 0.001 P< 0.001 P< 0.001 23225272353498 0.5 0.6 0.7 0.8 0.9 a b c 498 Total Categories 18,317 Total Samples Train:14410 | Test:3907 Dermnet498 ACC PanDerm SkinGPT-X MCC PanDerm SkinGPT-X Weighted F1 PanDerm SkinGPT-X Kappa PanDerm SkinGPT-X 0.580 0.634 0.639 0.557 0.573 0.625 0.577 0.632 Tinea Ringworm Candidiiasis & other Fungal Infections (1625) Seborrheic Keratoses & other Benign Tunmors (1714) Systemic Disease (758) Bullous Disease Photos (561) Light Diseases & Disorders of Pigentation (711) Vasculitis Photos (521) Cellulitis Impetigo & Bacterial Infections (361) Poison Ivy & other Cantact Dermatitis (325) Melanoma Skin Cancer Nevi & Moles (579) Herpes HPV & other STDs Photos (507) Psoriasis Lichen Planus & Related Diseases (1757) Warts Molluscum & other Viral Infections (1358) Atopic Dermatitis (613) Hair Loss Alopecia & other Hair Diseases (299) Nail Fungus & other Nail Diseases (1301) Actinic Keratosis Basal Call Carcinoma & other Malignant Lesions (1437) Eczema Photos (1544) Acne & Rosacea Photos (1152) Vascular Tumors (603) Scabies Lyme & other Infestations & Bites (539) Exanthems & Drug Eruptions (505) Lupus & other Connective Tissue Diseases (525) Urticaria Hives (265) Dermnet Dermnet272Dermnet353 Dermnet498 Dermnet225 Eczema & Dermatitis (2046) Fungal Infection (1067) Acne & Rosacea (787) Eczema & Atopic Dematitis (1161) Tinea (1039) Contact Dermatitis (788) Tinea (1039) Eczema (1007) Psoriasis (821) Fig. 3. a, Data distribution and hierarchical composition of the dataset series. This figure illustrates the incremental expansion of skin disease categories within the Dermnet dataset, scaling from Dermnet to the comprehensive Dermnet498. Each colored block represents a unique disease class. b, Comparative performance metrics on the Dermnet498 dataset. The bar chart presents ACC, MCC, Weighted F1, and Kappa scores for MedGemma, Hulu-Med, Qwen3-VL, PanDerm, and SkinGPT-X. Numerical values and error bars represent the mean and confidence intervals, with statistical significance between the top two models indicated by P -values. The radar chart displays the composition of Dermnet498. The legend specifies the top categories and their respective sample counts.c, Performance trends across varying category scales. The line graphs plot the changes in ACC, MCC, Weighted F1, and Kappa for the PanDerm and SkinGPT-X models as the number of disease categories increases. 7 pictures Lichen Planus and related diseases is decoupled into 27 distinct classes, including 17 Psoriasis subtypes and 6 vari- ants of Lichen Planus. Similarly, the Fungal Infection cluster is expanded into 24 specific categories, distinguishing site- specific infections such as CandidaAxillae and Candidasis Mouth. By shifting from 23 broad categories to 498 fine- grained categories, we establish a rigorous benchmark that mirrors the complexity of real-world dermatology. As shown in Fig. 3b, this hierarchy encompasses a di- verse distribution, ranging from common conditions to rare manifestations across 498 total categories. Detailed specifi- cations of the Dermnet498 hierarchy are provided in Section 4.2. As illustrated in Fig. 3c, we performed a comparative analysis against PanDerm by progressively increasing the diagnostic granularity from 23 to 498 categories. While the performance of both models naturally trends downward as the classification complexity intensifies, SkinGPT-X consis- tently maintains a robust margin over PanDerm. Notably, SkinGPT-X preserves its ACC, Weighted F1, and Cohen’s Kappa above the 60% threshold, whereas the baseline per- formance drops more precipitously. Asshowninthecomparativebenchmarksfor Dermnet498 (Fig. 3b), MedGemma, Qwen3-VL, and Hulu- Med, demonstrate significant performance degradation when faced with such highly specialized and challenging fine-grained classification tasks. In contrast, the fine-tuned foundation model PanDerm achieves an ACC of 58.0%, an MCC of 55.7%, a Weighted F1 of 57.3%, and a Kappa of 57.7%. Remarkably, our SkinGPT-X system demonstrates a significant performance leap across all dimensions (P < 0.001). Specifically, it reaches an ACC of 63.4% (+5.4%), an MCC of 63.9% (+8.2%), a Weighted F1 of 62.5% (+5.2%), and a Kappa of 63.2% (+5.5%). This superior performance on Dermnet498 underscores SkinGPT-X’s capacity for high- precision diagnosis. Its enhanced MCC and Weighted F1 suggest exceptional potential for the diagnosis of rare skin diseases, a domain where general LLMs lack sufficient spe- cialized prior knowledge. 2.4 SkinGPT-X supports rare skin disease diagnosis To further validate the robustness of SkinGPT-X in diagnos- ing rare and complex conditions where general LLMs often struggle due to the scarcity of training data, we constructed the Rare Skin Disease Dataset. As illustrated in Fig. 4a, the RSDD was rigorously curated by integrating diverse high- quality sources, including medical digital libraries, clinical centers, and specialized rare disease diagnosis and treat- ment guidelines. The final dataset comprises 564 samples across 8 distinct rare dermatological categories, with rep- resentative conditions such as Cutaneous Neuroendocrine Carcinoma (n = 115), Generalized Pustular Psoriasis (n = 110), and Behcet’s Disease (n = 81). As shown in the multi-dimensional performance com- parison (Fig.4b), SkinGPT-X achieves the highest diag- nostic efficacy across 5 evaluated metrics, including ACC, MCC, Macro F1, Weighted F1, and Kappa. The radar chart visually demonstrates that SkinGPT-X not only significantly outperforms general LLMs like MedGemma and Hulu-Med which exhibit extremely limited prior knowledge in this domain, but also surpasses the fine-tuned foundation model PanDerm. This superiority underscores the effectiveness of our proposed SkinGPT-X system in harmonizing visual features with specialized clinical knowledge. 2.5 EvoDerma-Mem is Indispensable for Evolving Clin- ical Reasoning and Expert-Level Knowledge Maturation EvoDerma-Mem enhances the diagnostic process by pro- viding evolved diagnostic guidelines and retrieved past cases. As illustrated in Fig. 4c, for the case of Blue Rubber Bleb Nevus Syndrome, by cross-referencing the evolved guidelines, the system recognized that while the classic color is blue, dark pigmentation is a known variant, especially when presented as multiple, smooth-surfaced lesions in a linear ar- rangement. This allowed the model to correctly identify a systemic syndrome rather than dismissing the findings as isolated melanocytic lesions. Similarly, in the diagnosis of Behcet’s Disease, the system analyzed a solitary, ery- thematous ulcerative lesion on the buccal mucosa. While such a finding could be non-specific, the retrieved past cases and guidelines from EvoDerma-Mem emphasized the importance of shallow, irregular borders with central necrotic debris and surrounding erythema. By comparing these visual findings against historical cases, the system was able to morphologically similar disorders such as Gorlin Syndrome an Blue Rubber Bleb Nevus Syndrome. EvoDerma-Mem emulates the iterative learning trajec- tory of a senior clinician. This continuous refinement is visualized in the Guidelines Evolution Timeline (Fig. 5b), where each node represents a specific iteration of a disease’s diagnostic guidelines. The size and color intensity of each bubble quantify the magnitude of descriptive refinement, relative to its previous version. A salient example is the evolution of Lichen Planus (LP) guidelines as shown in Fig. 5c. While the previous version was confined to the classic 6 P’s mnemonic (Pruritic, Planar, Purple, Polygonal, Papules, and Plaques), the current version has matured into a high-fidelity clinical manifest. By assimilating new cases, the system has extended its descriptive scope to include atypical anatomical sites (e.g., trunk and hands) and identified critical morphological mimics, such as dome- 8 MedGemma Blue Rubber Bleb Nevus Syndrome (73) Behcet's Disease (81) Cutaneous Neuroendocrine Carcinoma (115) Dermatofibrosarcoma Protuberans (59) Gorlin Syndrome (58) Generalized Pustular Psoriasis (110) Cryopyridine-associated periodic syndrome (22) Epithelioid Sarcoma (46) Kappa Macro F1 ACC Weighted F1 MCC Hulu-Med Qwen3-VL SkinGPT-X PanDerm Data Sources Wiley Online Library Stamford Skin Center Wikidoc Healthline ResearchGate The National Psoriasis Foundation Rare Disease Diagnosis and Treatment Guidelines Construction of Rare Disease Dataset (RDD) Disease Selection Criteria Categories and Sample Size of Final Dataset (n=564) Visual Findings: The image shows multiple discrete, well- defined, dark brown to black macules and papules on the upper chest and neck, set against a background of diffuse erythema. The lesions are smooth, lack scaling, crusting, ulceration, or telangiectasia, and have regular borders. The surrounding erythema suggests a separate inflammatory process. The lesions appear stable and non-infiltrative. Case Reviewing: While the description doesn't explicitly mention a 'blue' hue, the dark pigmentation and morphology are consistent with melanocytic lesions. The presence of multiple lesions, particularly the linear arrangement of three lesions below the clavicles, raises suspicion for a syndrome. The surrounding erythema is noted as a separate process. Comparing these findings with the provided diagnostic guidelines - Blue Rubber Bleb Nevus Syndrome: The description aligns well with the morphology of blue rubber bleb nevi (dark pigmented papules/macules, smooth surface, well-defined borders). Although the classic color is 'blue', dark brown/black pigmentation can occur, and the syndrome involves multiple lesions. Performance Comparison on Rare Disease Dataset a b c Primary Diagnosis: Blue Rubber Bleb Nevus Syndrome PrimaryDiagnosis: Behcet’s Disease. Visual Findings: The lesion depicted is a solitary, well- demarcated, erythematous ulcerative lesion located on the buccal mucosa. Morphologically, it presents as a shallow, irregularly bordered ulcer with a central area of necrotic debris and surrounding erythema. Case Reviewing: The visual description in detail an erythematous, ulcerative lesion on the buccal mucosa with irregular borders, central necrosis, and surrounding inflammation. This morphology strongly aligns with the typical presentation of oral ulcers in Behcet's disease. The diagnostic guidelines for Behcet Disease mention ulcers, asymmetry, lack of internal patterns, and inflammation, all consistent with the findings. Past cases also describe similar oral lesions consistent with Behcet’s Disease. The other options are less likely: Gorlin Syndrome typically involves basal cell carcinomas, not oral ulcers; Blue Rubber Bleb Nevus Syndrome involves specific blue papules and systemic features not described. Therefore, Behcet's disease is the most fitting diagnosis based on the visual findings, model prediction, and comparison with standard references and past cases. Fig. 4. a, Construction pipeline of the Rare Skin Disease Dataset. The dataset integrates diverse data sources, including medical libraries, clinical centers, and specialized guidelines to ensure high-quality labels for rare dermatological conditions. The RSDD dataset comprises 8 rare conditions, including Cutaneous Neuroendocrine Carcinoma (n = 115), Generalized Pustular Psoriasis (n = 110), and Behcet’s Disease (n = 81), among others. b, Multi-dimensional performance comparison on the RSDD (n = 564). The radar chart illustrates the diagnostic efficacy of SkinGPT-X versus MedGemma, Hulu-Med, Qwen3-VL, and PanDerm across five metrics: ACC, MCC, Macro F1, Weighted F1, and Kappa. c, Representative case studies of rare disease diagnosis and differential reasoning. EvoDerma-Mem enhances the capability of the system differentiating the correct diagnosis form rare candidate skin diseases shaped lesions and scaly plaques that frequently overlap with psoriasis or eczema. To evaluate the critical necessity of EvoDerma-Mem, we conducted ablation studies and clinical expert validations. As illustrated in Fig. 5a, ablation studies demonstrate that the integration of EvoDerma-Mem consistently yields per- formance gains across five metrics (ACC, BACC, Weighted F1, MCC, and Kappa, all p < 0.001).Specifically, on the Fitzpatrick17k dataset, the inclusion of evolved guidelines and historical case retrieval elevated the ACC from 68.4% to 70.7%, and the Weighted F1 from 68.4% to 72.5%, repre- senting a significant improvement in diagnostic precision. Similar trends were observed across other metrics, under- scoring the system’s enhanced reliability in dermatological diagnosis.Results on the Dermnet498 dataset further vali- date these findings, where the performance gains were even more pronounced. The full SkinGPT-X framework achieved superior performance compared to the version without memory, with an +10.1% improvement in ACC, a +11.0% improvement in Weighted F1, a +10.7% improvement in MCC, and a +10.1% improvement in Kappa. These results indicate that EvoDerma-Mem significantly enhances the framework’s ability to handle diverse pathological presen- tations, particularly for high-dimensional categories with limited training samples. To assess the practical clinical utility of the evolved knowledge, we invited two senior dermatologists to con- duct a blinded clinical review of 600 evolutionary cases. As illustrated in Fig. 5c The experts evaluated the EvoDerma- Mem system from three critical dimensions: 1) Rigorous- ness of Medical Logic; 2) Validity of Diagnostic guidelines Completeness; 3) Rationality of Clinical Manifestation Re- finement. EvoDerma-Mem achieves consistently excellent perfor- mance across both the Fitzpatrick-17k and Dermnet498 datasets. In the quantitative evaluations, the system main- tained exceptional scores on a 5-point scale across all three critical dimensions, achieving 4.3 and 4.7 for Medical Logic Rigorousness, 3.8 and 4.5 for Diagnostic guidelines Com- pleteness, and 4.0 and 4.8 for Clinical Manifestation Refine- ment on the Fitzpatrick-17k and Dermnet498 benchmarks, respectively. These results underscore that EvoDerma-Mem 9 Neutral Neutral Neutral Rational Partially Rational Less Rational Rational Partially Rational Less Rational Rational Partially Rational Less Rational Rationality of Clinical Manifestation Refinement Validity of Diagnostic Criteria Completeness Rigorousness of Medical Logic Irrational Irrational Irrational 012345 Dermnet498 0.4 0.5 0.6 0.7 0.8 0.4 0.5 0.6 0.7 0.8 Fitpatrick17kDermnet498 P < 0.001P < 0.001P < 0.001P < 0.001 P < 0.001P < 0.001P < 0.001P < 0.001 SkinGPT-X Without MemorySkinGPT-X Rationality of Clinical Manifestation Refinement Validity of Diagnostic Criteria Completeness Rigorousness of Medical Logic Previous Version Lichen planus(LP) remains a polymorphic inflammatory condition, classically described by the 6 P's (Pruritic, Planar, Purple, Polygonal Papules and Plaques), but its presentation is highly variable. Cutaneous manifestations span discrete, well-defined papules often favoring flexural areas or extremities , to confluent, irregular plaques sometimes mimicking eczematous or psoriatic processes. Color ranges from violaceous to erythematous/salmon, and even hyperpigmented brown/gray, Current Version Lichen planus(LP) presents a protean morphology, challenging clinicians despite its classic description. While the '6 P's' offer a mnemonic, the reality spans discrete, erythematous to violaceous papules often found on extremities like the hands or trunk, to confluent, scaly plaques potentially mimicking psoriasis or eczema. Lesions can be planar or dome-shaped, smooth or scaly, and color varies widely from pink/salmon to violaceous (implied by 'purple'), or even hyperpigmented brown/gray , especially in darker skin or during resolution. Physician Evaluation Form New Cases a b c d Results of Physician Validation Fitpatrick17k Version Index 0.7070.7250.7030.703 0.6340.6390.6250.632 Diagnosis Guidelines Evolution Timeline ACC Weighted F1MCCKappa 4.1 4.7 3.8 4.5 4.0 4.8 0.6840.6800.6840.680 0.5330.5150.5320.531 ACC Weighted F1MCCKappa Fig. 5. a, Performance comparison between SkinGPT-X and the baseline framework (without memory) on Fitzpatrick-17k and Dermnet498 datasets. Data are presented as mean metrics with 95% confidence intervals; statistical significance was determined via two-sided t-tests (p < 0.001). b, Spatiotemporal evolution trajectory of diagnostic guidelines. The bubble color intensity represent degree of knowledge refinement, where deeper shades indicate more substantial updates to the diagnostic guidelines. c, The standardized Physician Evaluation Form used for blinded clinical review, focusing on Rigorousness of Medical Logic, Validity of Diagnostic guidelines Completeness and Rationality of Clinical Manifestation Refinement. d, Results of physician validation across three critical dimensions. is not merely a supplementary module but the indispensable architectural foundation that enables SkinGPT-X to maintain professional rigor and descriptive depth in real-world clini- cal workflows. 3 DISCUSSION By emulating the cognitive pathways of clinical experts, SkinGPT-X effectively achieves a cognitive segregation of visual findings extraction from diagnostic decision-making. By retrieving standardized medical knowledge from the Skin Handbook at this critical junction, the system provides a rigorous clinical anchor to validate the Disease Diag- nosis Agent’s hypothesis for candidate diseases [77]. Be- yond static data, the EvoDerma-Mem mechanism allows the system to accumulate actual clinical experience, constantly refining its diagnostic guidelines through various historical cases. The system conducts a multi-dimensional alignment between current visual findings and historical cases. By synergizing book-based dermatological knowledge with dy- 10 namically evolved diagnostic guidelines, SkinGPT-X per- forms a rigorous validation of the initial hypothesis to iden- tify the most clinically consistent diagnosis. This workflow enhances both the traceability and interpretability of the rea- soning process. Through multi-agent synergistic verification and cross-logical associations, the system attains superior diagnostic reliability that avoids the ’black box’ limitations of monolithic large-scale models [44]. Simultaneously, the EvoDerma-Mem mechanism emulates the cognitive evolu- tion of senior dermatologists, which allows the system to supplement foundational medical knowledge with a more comprehensive layer of dermatological insights, ultimately facilitating highly granular diagnostics and significantly im- proving the efficiency of identifying atypical or rare clinical presentations. Memory has emerged as a pivotal frontier in the evo- lution of AI agents [87]. Within the domain of dermatol- ogy, most diagnostic systems currently rely on static RAG to provide agents with foundational medical knowledge. However, generalized medical knowledge often proves in- sufficient for dealing with complex clinical cases [43]. Fur- thermore, fine-tuning LLMs is frequently constrained by the limitations of few-shot learning; with sparse clinical data, these models often fail to capture specific, discriminative diagnostic features [88]. To address this, we designed the EvoDerma-Mem mechanism, which emulates the empirical synthesis process of experienced clinicians by continuously distilling standout details from novel cases to form evolv- ing diagnostic guidelines. During the diagnostic phase, EvoDerma-Mem provides agents with both visually similar historical cases and refined diagnostic guidelines. This dy- namically evolving agent memory offers superior heuristic guidance even in few-shot scenarios, while simultaneously ensuring a more transparent and traceable diagnostic ra- tionale by explicitly linking current findings to retrieved historical cases and their synthesized diagnostic guidelines [89]. The experimental results further illuminate the effec- tiveness of the EvoDerma-Mem mechanism, particularly in stabilizing the reasoning process as diagnostic complexity scales. As illustrated in Fig. 3c, while baseline models suffered significant performance degradation when tran- sitioning to fine-grained classification (up to 498 classes), SkinGPT-X maintained high diagnostic resilience, outper- forming specialized foundation models like PanDerm [66]. This advantage is even more pronounced in the RSDD, where the system’s ability to engage in guidelines-aligned inference allows it to transcend simple pattern matching, effectively addressing the challenges of few-shot clinical rea- soning. As evidenced in Fig. 5, both the ablation studies and the double-blind reviews conducted by senior dermatolo- gists substantiate that EvoDerma-Mem is an indispensable architectural component. The ablation of this mechanism results in a statistically significant degradation in diagnostic precision. The evolved guidelines generated by the system exhibit a high degree of feature comprehensiveness and inferential coherence, demonstrating a rigorous alignment with established professional clinical standards [90]. Despite the significant advancements demonstrated by SkinGPT-X, several intrinsic limitations regarding its archi- tectural complexity and clinical deployment warrant further discussion. First, the transition from monolithic models to a MAS framework introduces a necessary trade-off between reasoning depth and computational efficiency. The collab- orative orchestration required for visual feature extraction, EvoDerma-Mem retrieval, and iterative guideline synthesis results in prolonged processing times, which may pose challenges in high-throughput clinical settings where real- time feedback is critical [91]. Furthermore, the system’s diagnostic reliability is sensitive to the heterogeneity of im- age acquisition across different medical centers. Variations in sensor calibration, ambient lighting, and magnification levels can induce a distribution shift in the latent feature space, potentially leading to device-specific biases that de- grade the retrieval accuracy of the EvoDerma-Mem mech- anism [92]. In future research, we will focus on optimizing agent communication protocols and implementing robust domain-adaptation techniques to ensure both operational speed and cross-center generalizability. This study establishes a new paradigm for dermato- logical artificial intelligence by combining the self-evolving agent memory with systematic collaboration of multiple AI agents. The incorporation of the EvoDerma-Mem mecha- nism and guidelines-aligned reasoning strategies enhances diagnostic resilience, particularly in few-shot scenarios in- volving rare skin diseases, while simultaneously ensuring a more transparent and traceable diagnostic rationale. These findings suggest that the future of medical AI lies not merely in the expansion of static knowledge bases, but in the explicit emulation of clinical experience and the structural stabilization of deep-reasoning workflows. 4 METHODS 4.1 Metrics To comprehensively evaluate the performance of the pro- posed model, we employ five widely recognized metrics: ACC, Macro F1, Weighted F1, MCC, and Cohen’s Kappa. These metrics are derived from the confusion matrix com- ponents: True Positives (TP ), True Negatives (TN ), False Positives (FP ), and False Negatives (FN ). 11 ACC measures the proportion of correctly classified in- stances among the total number of cases, providing a gen- eral assessment of the model’s performance: ACC = TP + TN TP + TN + FP + FN (1) Macro F1 is the arithmetic mean of class-specific F1 values. It assigns equal weight to each class, ensuring that per- formance on minority classes is not overshadowed by the majority class in imbalanced datasets: Macro F 1 = 1 2 2TP 2TP + FP + FN + 2TN 2TN + FN + FP (2) Weighted F1 averages the F1-score of each class, weighted by their respective number of samples (N i ): Weighted F 1 = L X i=1 2· Precision i · Recall i Precision i + Recall i · N i N total (3) MCC is a reliable statistical rate that produces a high score only if the prediction performs well across all categories. To fit the column width, the formula is expressed as: MCC =(TP · TN − FP · FN )/ q (TP + FP )(TP + FN )(TN + FP )(TN + FN ) (4) Cohen’s Kappa measures the agreement between predicted and observed classifications while accounting for the possi- bility of agreement by chance: κ = p o − p e 1− p e (5) where p o is the observed agreement and p e is the expected agreement due to chance. 4.2 Dataset To evaluate the diagnostic efficacy and generalization of SkinGPT-X, we employed five benchmark datasets and cu- rated a specialized rare disease evaluation suite: Dermnet, one of the largest and most widely recognized public clinical image benchmarks for skin disease classifi- cation. This dataset comprises 18,856 images specifically cu- rated to represent the diversity of dermatological conditions encountered in clinical practice. The images are organized into 23 primary super-classes, covering a broad spectrum of common and complex pathologies. HAM10000, which stands as one of the most prominent and frequently cited public dermoscopic image benchmarks for the classification of pigmented skin lesions. This dataset comprises 10,015 high-quality dermoscopic images, metic- ulously curated to address the challenges of automated diagnosis in dermatology. The images are organized into 7 primary diagnostic categories, representing a significant spectrum of common benign and malignant pigmented pathologies. Fitzpatrick-17k, a large-scale clinical image benchmark specifically developed to address the critical need for skin tone diversity and algorithmic fairness in dermatological AI. This dataset comprises 16,577 clinical images curated to represent a wide variety of skin conditions across all six Fitzpatrick skin types (FST I to VI). Each image is meticulously annotated with both a skin condition label and its corresponding skin type, making it a pivotal resource for studying demographic bias. DDI31, a pioneering benchmark specifically designed to evaluate the fairness and performance of dermatology AI models across diverse skin tones. The dataset contains a total of 656 images, meticulously curated and patholog- ically confirmed from Stanford Clinics. In our study, we partitioned the dataset into a training set and a testing set using a 4 : 1 ratio to ensure a balanced evaluation. To accommodate a wide range of dermatological conditions, the dataset is categorized into 31 distinct diagnostic classes, encompassing both benign and malignant skin lesions. Dermnet225 is the most aggressively consolidated vari- ant among the sub-datasets, featuring a highly curated set of 225 classes. The mapping logic for Dermnet225 prioritizes clinical grouping to minimize noise from ambiguous sub- labels. We implemented a multi-tiered infection classifica- tion, where various fungal species were grouped under Fungal Infection and diverse bacterial presentations (e.g., Folliculitis, Impetigo, and Cellulitis) were unified into Bacterial Infection. Similarly, benign soft tissue growths and different types of viral warts were merged based on their biological commonalities. This dataset serves as a benchmark for eval- uating models on a consolidated yet diverse label space. Dermnet272 represents a more condensed hierarchical reorganization, targeting a classification space of 272 cate- gories. The construction of this dataset focused on reducing fragmented sub-classifications by applying deeper medical grouping logic. Beyond the standard class consolidation used in Dermnet353, we implemented an advanced merg- ing logic to handle varied spellings of complex conditions (e.g., merging multiple variants of Ichthyosis). Furthermore, various insect bites and stings were grouped into a unified Bites and Stings class, and miscellaneous nail disorders were aggregated to ensure each category maintained sufficient diagnostic representativeness. Dermnet353 is a refined intermediate-scale version of the dataset, designed to balance taxonomic breadth with 12 categorical distinctness. While the original Dermnet is struc- tured into 23 super-classes, we implemented a rule-based merging strategy to map the underlying fine-grained images into 353 clinically relevant categories. This process involved normalizing labels and consolidating sub-categories that share high morphological similarity, such as merging var- ious types of Contact Dermatitis (allergic, irritant, and phy- tophotodermatitis) into a single group. Additionally, para- sitic infestations like Scabies and Pediculosis were unified due to their overlapping clinical presentation in a computer vision context. Dermnet498, is a high-cardinality benchmark we con- structed by refining the original Dermnet dataset. While the original Dermnet is organized into 23 broad super-classes, it contains rich, hierarchical metadata where images are associated with specific clinical sub-categories. Leveraging this hierarchy, we reorganized the dataset by re-mapping the images from 23 coarse super-classes to 498 distinct fine-grained sub-classes. During the curation process, we performed a rigorous data cleaning step: images associated with sub-labels that lacked clinical significance or were diagnostically ambiguous were excluded. Consequently, the final Dermnet498, Dermnet353, Dermnet272, Dermnet225 datasets comprise 18,853 images, a slight reduction from the original count. This transformation shifts the objective from broad-category classification to a more challenging fine- grained recognition task within an extensive label space. The Dermnet498 dataset was partitioned into a training set of 14,856 images and a test set of 3,997 images, ensuring that both sets encompass the full spectrum of the 498 sub- classes. 4.3 The EvoDerma-Mem Mechanism in SkinGPT-X To simulate the clinical expertise accumulation process of a dermatologist, who refines diagnostic experience through continuous clinical practice. We developed a novel self- evolving agent memory name EvoDerma-Mem for dermato- logical diagnosis, as illustrated in Fig. 1(b). This mechanism comprises a structured database and a knowledge evolution process. 4.3.1 Case Representation and agent memory Storage Initially, each historical sample S i in the repository is processed through a pre-trained Feature Extractor Φ. This module encodes the high-dimensional semantic features of the clinical image I i into a compact embedding codez i : z i = Φ(I i ),z i ∈R d (6) Each memory entry M i is stored as a linked triplet within a graph database, formally defined as: M i =⟨z i ,K i ,D i ⟩(7) where K i represents the textual Key Findings and D i denotes the corresponding expert diagnosis. During inference, the system retrieves the most relevant historical cases by calcu- lating the cosine similarity between the query embedding z q and the stored embeddings: Score(M q ,M i ) = z q ·z i ∥z q ∥z i ∥ (8) 4.3.2 Knowledge Synthesis and Guideline Evolution The knowledge evolution process leverages a specialized Summarization Agent A to synthesize collective diag- nosis guidelines. For a specific disease category C , the agent analyzes the set of all associated findings K C = K 1 ,K 2 ,...,K n to generate the initial diagnostic guide- lines G 0 C : G 0 C =A(K C )(9) The system implements a dynamic evolution process triggered by an accumulation threshold N thresh . When the number of new samples ∆N ≥ N thresh , the Summariza- tion Agent updates the current guidelines G t C to the next iteration G t+1 C by integrating the latest findingsK new : G t+1 C =A(G t C ⊕K new )(10) where ⊕ represents the knowledge integration operation that resolves contradictions and incorporates novel mor- phological evidence. This closed-loop evolution ensures that SkinGPT-X continuously refines its clinical reasoning logic, achieving a progressive maturation of diagnostic intelli- gence as the longitudinal dataset grows. 4.4 Multi-Agent Collaborative Diagnosis Framework To achieve the transparent and trustworthy dermatological reasoning, SkinGPT-X implements a collaborative frame- work comprising multiple specialized agents as shown in Algorithm 1. Algorithm 2 provides the comprehensive com- putational procedures for both the autonomous agents and the knowledge-retrieval functions within the framework. 13 Algorithm 1. SkinGPT-X Inference Process Require: Input patient image I , Skin handbook B, Agent memory M (containing historical cases K hist and evolved guidelinesG). Ensure: Final validated diagnostic report R ∗ and optimal diagnosis D ∗ . 1:z I ← VisionEncoder(I) ▷ Generate visual embedding 2: P vis ← VisualAgent(Prompt morph ,I)▷ Objective morphological 3: D pre ← Pre-DiagAgent(I) ▷ Top-5 candidate diseases 4: for each d i ∈D pre do 5: K prior,i ← Retrieve(B,d i )▷ Fetch textbook standards 6: G evo ← S d i ∈D pre G[d i ] ▷ Self-evolved guidelines 7: K hist ← Query(M,z I ) ▷ Retrieve top-5 similar cases 8: X rev ←P vis ,D pre ,K prior ▷ Construct evidence space 9: D ∗ ←A rev (X rev ,K hist ,G evo ) ▷ Cross-verification, details are listed in Section 4.4.3 10: R ∗ ← (D ∗ ,P vis ) 11: return Validated ReportR ∗ 4.4.1 Independent Visual and Diagnostic Agents The reasoning pipeline begins with the Vision Agent, pow- ered by the Qwen3-VL model. Rather than providing a definitive diagnosis, this agent is restricted to the objective characterization of pathological features from the input im- agery. It generates a structured description of the lesion’s morphology, pigmentation patterns, and boundary condi- tions, establishing a factual foundation for downstream reasoning. Subsequently, the Disease Diagnosis Agent utilizes a scalable architecture designed for modular integration. This module can embed one or more dermatology-specific foun- dation models as pluggable components. The agent syn- thesizes the visual evidence to output the top-5 candidate diagnoses D = d 1 ,...,d 5 , where each diagnosis d i is associated with a corresponding confidence score P (d i ) for i∈1,..., 5. 4.4.2 Knowledge-Augmented Reasoning via RAG To reinforce diagnostic rigor, we integrate a RAG module supported by a custom-built Skin Handbook. This knowl- edge base B was constructed by extracting domain-specific expertise from the Oxford Handbook of Medical Dermatol- ogy [93], covering the structure and function of the skin, standardized diagnostic standards, and distinguish features of diverse skin conditions. The RAG module retrieves specific medical standards for each candidate disease: K prior = Retrieve(B,d i )(11) Algorithm 2. Details of the computational procedures for agents and the knowledge-retrieval functions Require: Input patient image I ; Skin handbook B; Agent memory M (containing historical cases K hist and evolved guidelinesG evo ); Function Embedding(I): 1:z I = PanDerm frozen (I) 2:returnz I ∈R d .▷ Visual Embeddings ofI Function VisualAgent(Prompt morph ,I): 3: P vis = LLM Qwen3-VL (Prompt morph ⊕I) 4:returnP vis ;▷ Morphological descriptions ofI Function Pre-DiagnosisAgent(I): 5: P(d|I) = Softmax(PanDerm tuned (I)) 6: D pre =(d i ,p i )| rank(P (d i |I))≤ 5 7:returnD pre ▷ Diagnostic candidate set, where d i is the disease label and p i is the confidence score; Function Retrieve(B,d i ): 8: K prior = arg max k∈B cos(Emb text (d i ), Emb text (k)) 9:returnK prior ▷ Medical standards from textbook Function Query(M,z I ): 10: J ∗ = TopK j z I ·z j ∥z I ∥z j ∥ , s.t. (z j ,d j )∈M,K = 5 11: K hist =(d j ,P j ) j∈J ∗ 12:returnK hist ▷ Similar historical cases from agent momery 4.4.3 The Enhanced Case Reviewing based on Self- Evolving Agent Memory To achieve high-fidelity diagnostic conclusions, we design a Case Reviewing agent A rev powered by the Qwen3-A30 [85]. This module functions as a clinical reviewer, perform- ing a traceable review of multi-source evidence to emulate the rigorous synthesis process of senior dermatologists. First, diverse clinical information mentioned above is collected and organized to provide a unified basis for multi- dimensional cross-verification. Formally, for a given sample q, the input spaceX rev is defined as: X rev =P vis ,D pre ,K prior (12) where: • P vis represents the Multimodal Perception, includ- ing visual pathological descriptions and key find- ings. • D pre = (d i ,p i ) 5 i=1 is the list of Candidate Diag- noses with their respective confidence scores. 14 • K prior = Retrieve(B,d i ) denotes the Expert Knowl- edge Alignment, retrieving standardized standards from the Skin HandbookB. Second, the agent evaluates the inferential coherence across multiple information dimensions to derive the op- timal diagnosis D ∗ . Unlike conventional models that rely on frozen parameters θ, our system leverages the in-context learning capabilities of the LLM to orchestrate foundational medical knowledge with experiential clinical insights: D ∗ ←A val (X rev ,K hist ,G evo )(13) whereK hist andG evo are retrieved from the agent memory. K hist represents the Historical Cases, containing the top-5 similar cases’ visual key findings and their diagnosis. G evo signifies the self-evolved diagnostic guidelines refined by the Summarize Agent. By incorporating G evo , the system transcends simple instance-based retrieval, leveraging high- level clinical abstractions synthesized from collective his- torical encounters. This mechanism ensures that the final output is not a mere statistical prediction, but a clinically validated conclusion grounded in the synergistic verifica- tion of textbook standards and empirical precedents. The reasoning process is operationalized through a five- stage Review-and-Synthesis protocol: •Stage 1: Visual Feature Validation. The agent first performs an independent review of the initial per- ceptionP vis . It identifies and rectifies potential omis- sions in morphological descriptors to ensure a high- fidelity visual foundation. •Stage 2: Canonical Guidelines Cross-Check. The validated findings are cross-referenced againstK prior and G evo . The agent assesses whether the observed lesion meets the discriminative features defined in standardized clinical guidelines. •Stage 3: Empirical Evidence Alignment. By analyz- ingK hist , the agent evaluates the consistency between the current case and confirmed historical precedents, providing empirical grounding that transcends static model probabilities. •Stage 4: Conflict Resolution & Systematic Synthe- sis. In cases when statistical predictionsD pre conflict with clinical guidelines, the agent prioritizes author- itative standards, identifying if the visual model is influenced by non-diagnostic artifacts. •Stage 5: Final Diagnostic Determination. The syn- thesized evidence is mapped to a final decision D ∗ , derived from the candidate setD pre . 4.4.4 The curation of Rare Skin Disease Dataset (RSDD) To evaluate the generalization capabilities and diagnostic robustness of SkinGPT-X within the long-tail distribution of dermatology, we curated the RSDD, a novel, high-quality benchmark specifically focused on rare skin conditions. As illustrated in Fig. 4 a, the construction of RSDD followed a rigorous three-step pipeline: •1) Disease Selection and Taxonomy Alignment: The curation of this dataset was conducted in accordance with the Rare Disease Diagnosis and Treatment Guide- lines published by the National Health Commission of China [94]. From this official compendium, 12 dermatological conditions were initially identified. To ensure the integrity of the evaluation and prevent data leakage, we cross-referenced these with our pri- mary training corpus. Overlapping conditions (e.g., Melanoma) were systematically excluded. •2) Data Acquisition from Diverse Sources: To ensure high-fidelity and authoritative clinical repre- sentation, we aggregated samples from multiple rep- utable medical platforms and repositories, including Wiley Online Library (https://onlinelibrary.wiley.com/), StamfordSkinCenter(https://stamfordskin.com/), Wikidoc(https://w.wikidoc.org/),Health- line(https://w.healthline.com/),Healthline (https://w.healthline.com/),ResearchGate (https://w.researchgate.net/),andTheNational Psoriasis Foundation (https://w.psoriasis.org/). •3) Final Cohort and Experimental Design: The final dataset (n = 564) comprises 8 distinct rare condi- tions, including Cutaneous Neuroendocrine Carcinoma (115), Generalized Pustular Psoriasis (110), Behcet’s Dis- ease (81), and Blue Rubber Bleb Nevus Syndrome (73), Epithelioid Sarcoma (46), Cryopyridine-associated peri- odic syndrome (22), Dermatofibrosarcoma Protuberans (59), Gorlin Syndrome ((58)). To simulate the data- scarce reality of rare disease clinical practice, we implemented a 1 : 2 train-test split. This constrained distribution is designed to stress-test the model’s few-shot learning efficiency and its ability to gener- alize from limited clinical presentations compared to general-purpose LLMs. 4.5 Implementation Details During the finetuning process of Disease Diagnosis Agent, the max number of epochs is fixed to 30, the warmup epoch is fixed to 10, the learning rate is set to 5e-4 and the batch size is fixed to 128. For other agents in SkinGPT-X, we use Qwen3 series models as the foundation, the maximum token number is fixed to 4096 and the temperature is set to 0.3. 15 To deploy SkinGPT-X entirely locally, a Linux system(e.g. Ubuntu 18.04) is mandatory. For acceleration, we recom- mend using 8 4090 GPUs with 24GB of memory. SkinGPT-X is developed using Python 3.10, PyTorch 1.10.2 and CUDA 12.5. 5 ACKNOWLEDGEMENTS Funding: This work is supported by The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), under Award No. UDF01004172. Competing Interests: The authors have declared no com- peting interests. Author Contribution Statements: Z.C., Y.S. and J.Z. con- ceived of the presented idea. Z.C., F.W., H.C. and W.D. de- signed the computational framework and analysed the data. Z.C., Y.X., L.S. and J.Z. conducted the clinical evaluation. J.Z. supervised the findings of this work. Z.C., Y.S., Z.W. and J.Z. took the lead in writing the manuscript and supple- mentary information. All authors discussed the results and contributed to the final manuscript. Data availability: The data that support the findings of this study are divided into two groups: shared data and restricted data. Shared data include the HAM10000, DDI, Fitzpatrick-17k, and Dermnet datasets. The HAM10000 dataset is accessible via the Harvard Dataverse at https: //doi.org/10.7910/DVN/DBW86T; the DDI dataset can be accessed through its official project repository at https://ddi-dataset.github.io/; the Fitzpatrick-17k dataset is available via its GitHub repository at https://github. com/mravuri/fitzpatrick17k; and the Dermnet dataset can be accessed at https://w.kaggle.com/datasets/ shubhamgoel27/dermnet. Restricted data, including private clinical records and certain proprietary benchmarks, are not publicly available due to patient privacy and institutional data-sharing agreements. The restricted in-house skin dis- ease images used in this study are not publicly available due to restrictions in the data-sharing agreement. Code availability: To promote academic exchanges, un- der the framework of data and privacy security, the code proposed by SkinGPT-X is publicly available at https:// github.com/healme-225040511/Skingpt X. In the case of non-commercial use, researchers can sign the license pro- vided in the above link and contact J.Z. or Z.C. REFERENCES [1]R. J. Hay, N. E. Johns, H. C. Williams, I. W. Bolliger, R. P. Dellavalle, D. J. Margolis, R. Marks, L. Naldi, M. A. Weinstock, S. K. Wulf et al., “The global burden of skin disease in 2010: an analysis of the prevalence and impact of skin conditions,” Journal of investigative dermatology, vol. 134, no. 6, p. 1527–1534, 2014. [2]N. Salari, P. Heidarian, A. Hosseinian-Far, F. Babajani, and M. Mo- hammadi, “Global prevalence of anxiety, depression, and stress among patients with skin diseases: A systematic review and meta- analysis,” Journal of Prevention, vol. 45, no. 4, p. 611–649, 2024. [3]F. Gorouhi et al., “Global, regional, and national burden of skin and subcutaneous diseases, 1990–2021: a systematic analysis for the global burden of disease study 2021,” The Lancet Rheumatology, 2024. [4]D. Seth, K. Cheldize, D. Brown, and E. E. Freeman, “Global burden of skin disease: inequities and innovations,” Current dermatology reports, vol. 6, no. 3, p. 204–210, 2017. [5]E. J. Emanuel et al., “Measuring the burden of health care utiliza- tion,” JAMA, vol. 324, no. 13, p. 1273–1274, 2020. [6]E. M. Warshaw et al., “Teledermatology for diagnosis and man- agement of skin conditions: a systematic review,” Journal of the American Academy of Dermatology, vol. 64, no. 4, p. 759–772, 2011. [7]J. J. Lee and J. C. English I, “Teledermatology: a review and update,” American journal of clinical dermatology, vol. 19, no. 2, p. 253–260, 2018. [8]R. V. Tuckson, M. Edmunds, and M. L. Hodgkins, “Advanced health information technology and electronic health records in the era of telemedicine,” New England Journal of Medicine, vol. 377, no. 16, p. 1585–1592, 2017. [9]J. D. Whited, “Teledermatology: current status and future direc- tions,” American Journal of Clinical Dermatology, vol. 2, no. 2, p. 59–64, 2001. [10] B. A. Moy et al., “The virtual visit triage: Challenges and successes of teledermatology during the COVID-19 pandemic,” Journal of the American Academy of Dermatology, vol. 84, no. 4, p. 1117–1118, 2021. [11] A. J. Thirunavukarasu et al., “Large language models in medicine,” Nature Medicine, vol. 29, no. 8, p. 1930–1940, 2023. [12] A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, and S. Thrun, “Dermatologist-level classification of skin cancer with deep neural networks,” Nature, vol. 542, no. 7639, p. 115– 118, 2017. [13] E. J. Topol, “High-performance medicine: the convergence of hu- man and artificial intelligence,” Nature medicine, vol. 25, no. 1, p. 44–56, 2019. [14] A. Gomolin, E. Netchiporouk, R. Gniadecki, and I. V. Litvinov, “Artificial intelligence applications in dermatology: where do we stand?” Frontiers in medicine, vol. 7, p. 100, 2020. [15] S. Chan, V. Reddy, B. Myers, Q. Thibodeaux, N. Brownstone, and W. Liao, “Machine learning in dermatology: Current applications, opportunities, and limitations,” Dermatology and Therapy, vol. 10, no. 3, p. 365–386, 2020. [16] S. S. Noronha, M. A. Mehta, D. Garg, K. Kotecha, and A. Abra- ham, “Deep learning-based dermatological condition detection: A systematic review with recent methods, datasets, challenges, and future directions,” IEEE Access, vol. 11, p. 13 988–14 015, 2023. [17] J. Dunn et al., “Deep learning for melanoma detection and seg- mentation,” Dermatology Practical & Conceptual, 2020. [18] M. A. Al-masni et al., “Multiple skin lesions classification using deep convolutional neural network,” IEEE Access, 2020. [19] A. Esteva et al., “Dermatologist-level classification of skin cancer with deep neural networks,” Nature, 2017. [20] A. Panthakkan et al., “A deep learning based intelligent clinical decision support system for skin lesion detection,” Ecological Infor- matics, 2022. [21] W. Gouda et al., “Skin cancer classification using resnet, inception and xception models,” Machine Learning and Knowledge Extraction, 2022. 16 [22] S. K. Bandyopadhyay, P. Bose, A. Bhaumik, and S. Poddar, “Ma- chine learning and deep learning integration for skin diseases pre- diction,” International Journal of Engineering Trends and Technology, vol. 70, no. 2, p. 11–18, 2022. [23] S. Jha and A. K. Mehta, “An evolutionary algorithm based feature selection and fuzzy rule reduction technique for the prediction of skin cancer,” Concurrency and Computation: Practice and Experience, vol. 34, no. 5, p. e6694, 2022. [24] S. S. Han, M. S. Kim, W. Lim, G. H. Park, I. Park, and S. E. Chang, “Classification of the clinical images for benign and malignant cutaneous tumors using a deep learning algorithm,” Journal of Investigative Dermatology, vol. 138, no. 7, p. 1529–1538, 2018. [25] B. Erkol, R. H. Moss, R. J. Stanley, W. V. Stoecker, and E. Hvatum, “Automatic lesion boundary detection in dermoscopy images using gradient vector flow snakes,” Skin Research and Technology, vol. 11, no. 1, p. 17–26, 2005. [26] T. Mazhar, I. Haq, A. Ditta, S. A. H. Mohsan, F. Rehman, I. Zafar, J. A. Gansau, and L. P. W. Goh, “The role of machine learning and deep learning approaches for the detection of skin cancer,” Healthcare, vol. 11, no. 3, p. 415, 2023. [27] A. Lembhe, P. Motarwar, R. Patil, and S. Elias, “Enhancement in skin cancer detection using image super resolution and convo- lutional neural network,” Procedia Computer Science, vol. 218, p. 164–173, 2023. [28] V. Venugopal, N. I. Raj, M. K. Nath, and N. Stephen, “A deep neu- ral network using modified efficientnet for skin cancer detection in dermoscopic images,” Decision Analytics Journal, vol. 8, p. 100278, 2023. [29] S. Jha and A. K. Mehta, “A hybrid approach using the fuzzy logic system and the modified genetic algorithm for prediction of skin cancer,” Neural Processing Letters, vol. 54, no. 2, p. 751–784, 2022. [30] A. H. Haenssle, C. Fink, R. Schneiderbauer, F. Toberer, T. Buhl, A. Blum, and A. Kalloo, “Man against machine: Diagnostic per- formance of a deep learning convolutional neural network for dermoscopic melanoma recognition in comparison to 58 derma- tologists,” Annals of Oncology, vol. 29, no. 8, p. 1836–1842, 2018. [31] S. M. Jaisakthi, P. Mirunalini, C. Aravindan, and R. Appavu, “Clas- sification of skin cancer from dermoscopic images using deep neural network architectures,” Multimedia Tools and Applications, vol. 82, no. 10, p. 15 763–15 778, 2023. [32] M. Zafar, M. I. Sharif, S. Kadry, S. A. C. Bukhari, and H. T. Rauf, “Skin lesion analysis and cancer detection based on machine/deep learning techniques: A comprehensive survey,” Life, vol. 13, no. 1, p. 146, 2023. [33] S. Gilmore, R. Hofmann-Wellenhof, and H. P. Soyer, “A support vector machine for decision support in melanoma recognition,” Experimental Dermatology, vol. 19, no. 9, p. 830–835, 2010. [34] R. Marks, “Epidemiology of melanoma,” Clinical and Experimental Dermatology, vol. 25, no. 6, p. 459–463, 2000. [35] N. Melarkode, K. Srinivasan, S. M. Qaisar, and P. Plawiak, “Ai- powered diagnosis of skin cancer: A contemporary review, open challenges and future research directions,” Cancers, vol. 15, no. 4, p. 1183, 2023. [36] S. Q. Gilani, T. Syed, M. Umair, and O. Marques, “Skin cancer classification using deep spiking neural network,” Journal of Digital Imaging, vol. 36, no. 3, p. 1137–1147, 2023. [37] N. Priyadharshini, B. Hemalatha, and C. Sureshkumar, “A novel hybrid extreme learning machine and teaching–learning-based- optimization algorithm for skin cancer detection,” Healthcare Ana- lytics, vol. 3, p. 100161, 2023. [38] M. A. Kadampur and S. Al Riyaee, “Skin cancer detection: Ap- plying a deep learning based model driven architecture in the cloud for classifying dermal cell images,” Informatics in Medicine Unlocked, vol. 18, p. 100282, 2020. [39] H. K. Gajera, D. R. Nayak, and M. A. Zaveri, “A comprehensive analysis of dermoscopy images for melanoma detection via deep cnn features,” Biomedical Signal Processing and Control, vol. 79, p. 104186, 2023. [40] P. Thapar, M. Rakhra, G. Cazzato, and M. S. Hossain, “[retracted] a novel hybrid deep learning approach for skin lesion segmentation and classification,” Journal of Healthcare Engineering, vol. 2022, no. 1, p. 1709842, 2022. [41] A. G. Pacheco and R. A. Krohling, “The impact of patient clinical information on automated skin cancer detection,” Computers in Biology and Medicine, vol. 116, p. 103545, 2020. [42] M. Tahir, A. Naeem, H. Malik, J. Tanveer, R. A. Naqvi, and S.- W. Lee, “Dscc net: Multi-classification deep learning models for diagnosing of skin cancer using dermoscopic images,” Cancers, vol. 15, no. 7, p. 2179, 2023. [43] S. Jabbour and et al., “Deep learning in dermatology: Challenges on the road to clinical implementation,” The Lancet Digital Health, 2023. [44] M. Moor, O. Banerjee, Z. S. H. Abad, H. M. Krumholz, J. Leskovec, T. Loh et al., “Foundation models for generalist medical artificial intelligence,” Nature, vol. 616, no. 7956, p. 259–265, 2023. [45] J. Achiam et al., “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774, 2023. [46] K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl et al., “Large language models encode clinical knowledge,” Nature, vol. 620, no. 7972, p. 172–180, 2023. [47] S. Ji et al., “Domain-specific integration of large language models: A survey,” arXiv preprint arXiv:2308.10020, 2023. [48] A. De, A. Sarda, S. Gupta, and S. Das, “Use of artificial intelligence in dermatology,” Indian Journal of Dermatology, vol. 65, p. 352, 2020. [49] K. Cirone, M. Akrout, L. Abid, and A. Oakley, “Assessing the utility of multimodal large language models (gpt-4 vision and large language and vision assistant) in identifying melanoma across different skin tones,” JMIR Dermatology, vol. 7, p. e55508, 2024. [50] X. Liu, C. Duan, M. K. Kim et al., “Claude 3 opus and chatgpt with gpt-4 in dermoscopic image analysis for melanoma diagno- sis: Comparative performance analysis,” JMIR Medical Informatics, vol. 12, p. e59273, 2024. [51] N. Shifai, R. van Doorn, J. Malvehy, and T. E. Sangers, “Can chatgpt vision diagnose melanoma? an exploratory diagnostic accuracy study,” Journal of the American Academy of Dermatology, vol. 90, p. 1057–1059, 2024. [52] T. Laohawetwanit, C. Namboonlue, and S. Apornvirat, “Accuracy of gpt-4 in histopathological image detection and classification of colorectal adenomas,” Journal of Clinical Pathology, vol. 77, no. 5, p. 304, 2024. [53] C. W. Rundle, M. D. Szeto, C. L. Presley et al., “Analysis of chatgpt generated differential diagnoses in response to physical exam findings for benign and malignant cutaneous neoplasms,” Journal of the American Academy of Dermatology, vol. 90, p. 615–616, 2024. [54] M. Sievert, M. Aubreville, S. K. Mueller et al., “Diagnosis of malig- nancy in oropharyngeal confocal laser endomicroscopy using gpt 4.0 with vision,” European Archives of Oto-Rhino-Laryngology, vol. 281, p. 2115–2122, 2024. [55] J. A. Omiye, H. Gui, S. J. Rezaei et al., “Large language models in medicine: The potentials and pitfalls: A narrative review,” Annals of Internal Medicine, vol. 177, p. 210–220, 2024. [56] A. I. Chen, L. K. Ferris, V. E. Nambudiri, and E. W. Piette, “Chatrx: Chatgpt’s potential to educate patients on medication adverse effects,” Journal of the American Academy of Dermatology, vol. 90, p. 669–670, 2024. [57] A. Breneman, M. H. Trager, E. R. Gordon, and F. H. Samie, “Read- ability rescue: Large language models may improve readability of 17 patient education materials,” Archives of Dermatological Research, vol. 316, p. 669, 2024. [58] M. A. Robinson, M. Belzberg, S. Thakker et al., “Assessing the accuracy, usefulness, and readability of artificial-intelligence- generated responses to common dermatologic surgery questions for patient education,” Journal of the American Academy of Derma- tology, vol. 90, p. 1078–1080, 2024. [59] K. C. Lauck, S. W. Cho, M. DaCunha et al., “The utility of artificial intelligence platforms for patient-generated questions in mohs micrographic surgery: A multi-national, blinded expert panel eval- uation,” International Journal of Dermatology, 2024. [60] G. M. El Saadawi, E. Tseytlin, E. Legowski et al., “A natu- ral language intelligent tutoring system for training patholo- gists—implementation and evaluation,” Advances in Health Sci- ences Education, vol. 13, p. 709, 2007. [61] Y. Zhang, R. Chen, D. Nguyen et al., “Assessing the ability of an artificial intelligence chatbot to translate dermatopathology reports into patient-friendly language: A cross-sectional study,” Journal of the American Academy of Dermatology, vol. 90, p. 397– 399, 2024. [62] K. Liopyris, S. Gregoriou, J. Dias, and A. J. Stratigos, “Artificial intelligence in dermatology: Challenges and perspectives,” Der- matologic Therapy, vol. 12, p. 2637–2651, 2022. [63] S.-C. Huang, L. Shen, I. Zhelezniak et al., “Visual language foun- dation models for the interpretation of medical images,” Nature Medicine, 2023. [64] T. Tu et al., “Towards generalist biomedical ai,” NEJM AI, vol. 1, no. 3, p. AIoa2300109, 2024. [65] M. Lewandowski, J. Kropidłowska, A. Kvinen, and W. Bara ́ nska- Rybak, “A systemic review of large language models and their implications in dermatology,” Australasian Journal of Dermatology, 2025. [66] J. Zhou and et al., “Skingpt-4: An interactive system for skin disease diagnosis that can read images,” arXiv preprint arXiv:2304.10674, 2023. [67] H. Zhang et al., “Huatuogpt, towards taming language models to be a chinese doctor,” arXiv preprint arXiv:2305.15075, 2023. [68] Y. Shen, Z. Chen, Y. He, Y. Xu, S. Zhang, L. Sun, Z. Wang, Y. Zhu, Y. Yang, J. Qian, Z. Wang, X. Zhang, W. Liu, Z. Ge, T. Lu, S. Yan, and J. Zhou, “Trustworthy and fair skingpt-r1 for democratizing dermatological reasoning across diverse ethnicities,” 2026. [Online]. Available: https://arxiv.org/abs/2511.15242 [69] F. Greco et al., “Limitations of large language models in healthcare: A systematic review,” npj Digital Medicine, vol. 6, no. 1, p. 1–13, 2023. [70] M. Ghassemi, L. Oakden-Rayner, and A. L. Beam, “The false hope of current explainable ai in health care,” The Lancet Digital Health, vol. 3, no. 11, p. e745–e750, 2021. [71] X. Wang, J. Wei, D. Schuurmans, L. Quoc, E. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Unleashing the emergent cognitive capacity of llms via joint optimization of masking and prompting,” arXiv preprint arXiv:2309.02427, 2023. [72] C.-M. Chan, W. Chen, Y. Su, J. Yu, W. Wei, Z. Zhang, Y. Lin, N. Wang, H.-N. Loh, Z. Liu et al., “Chateval: Towards better llm-based evaluators through multi-agent debate,” arXiv preprint arXiv:2308.07201, 2023. [73] G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “Camel: Communicative agents for” mind” exploration of large language model societies,” arXiv preprint arXiv:2303.17778, 2023. [74] X. Tang et al., “Medagents: Large language models as collaborators for zero-shot medical reasoning,” arXiv preprint arXiv:2311.10537, 2023. [75] T. Khot et al., “Decomposed prompting: A modular approach for solving complex tasks,” arXiv preprint arXiv:2210.02406, 2022. [76] S. S. Han and et al., “Classification of the clinical images for benign and malignant cutaneous tumors using a deep learning algorithm,” Journal of Investigative Dermatology, 2020. [77] C. Zakka et al., “Aloha: A clinical reasoning agent for healthcare,” Nature Medicine, 2024. [78] J. Soong and et al., “Improving medical reasoning with multi-agent retrieval-augmented generation,” arXiv preprint arXiv:2401.x, 2024. [79] S. Yan, Z. Yu, C. Primiero, C. Vico-Alonso, Z. Wang, L. Yang, P. Tschandl, M. Hu, L. Ju, G. Tan et al., “A multimodal vision foundation model for clinical dermatology,” Nature Medicine, p. 1–12, 2025. [80] DermNet, “Dermnet nz: The dermatology resource,” 2021. [Online]. Available: https://dermnetnz.org/ [81] P. Tschandl, C. Rosendahl, and H. Kittler, “The HAM10000 dataset, a large collection of multi-source dermatoscopic images of com- mon pigmented skin lesions,” Scientific Data, vol. 5, no. 1, p. 1–9, 2018. [82] R. Daneshjou, K. Vodrahalli, R. A. Novoa et al., “Disparities in dermatology AI performance on a diverse, curated clinical image set,” Science Advances, vol. 8, no. 32, p. eabq6147, 2022. [83] M. Groh, C. Harris, L. Soenksen, F. Lau, R. Han, A. Kim, A. Koochek, and O. Badri, “Evaluating deep neural networks trained on clinical images in dermatology with the Fitzpatrick 17k dataset,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, p. 1820–1828. [84] A. Sellergren, S. Kazemzadeh, T. Jaroensri, A. Kiraly, M. Traverse, T. Kohlberger, S. Xu, F. Jamil, C. Hughes, C. Lau et al., “Medgemma technical report,” arXiv preprint arXiv:2507.05201, 2025. [85] Q. Team, “Qwen3 technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2505.09388 [86] S. Jiang, Y. Wang, S. Song, T. Hu, C. Zhou, B. Pu, Y. Zhang, Z. Yang, Y. Feng, J. T. Zhou, J. Hao, Z. Chen, R. Wu, T. Tang, J. Lv, H. Xu, H. Wang, J. Xiao, B. Feng, F. Zhu, K. Li, W. Xie, J. Sun, J. Wu, and Z. Liu, “Hulu-med: A transparent generalist model towards holistic medical vision-language understanding,” 2025. [Online]. Available: https://arxiv.org/abs/2510.08668 [87] L. Wang and et al., “A survey on large language model based autonomous agents,” Frontiers of Computer Science, 2024. [88] O. Vinyals and et al., “Matching networks for one shot learning,” in Advances in Neural Information Processing Systems, 2016. [89] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. K ̈ uttler, M. Lewis, W.-t. Yih, T. Rockt ̈ aschel et al., “Retrieval-aug ented generation for knowledge-intensive nlp tasks,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, p. 9459–9474. [90] C. J. Kelly and et al., “Key challenges for delivering clinical impact with artificial intelligence,” BMC Medicine, 2019. [91] E. Kasneci, K. Seßler, S. K ̈ uchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. G ̈ unnemann, E. H ̈ ullermeier et al., “Chatgpt for good? on Opportunities and Challenges of Large Language Models for Education and Academic Research,” Learning and Individual Differences, vol. 103, p. 102274, 2023. [92] L. N. Guo and et al., “Algorithmic bias in skin cancer applica- tions,” JAMA Dermatology, 2021. [93] S. Griffiths and S. Burge, Oxford Handbook of Medical Dermatology, 2nd ed.Oxford University Press, 2020. [Online]. Available: https://academic.oup.com/book/26194 [94] General Office of the National Health Commission of the Peo- ple’s Republic of China, “Notice on issuing the diagnosis and treatment guidelines for 86 rare diseases, including achondropla- sia (2025 edition),” https://w.nhc.gov.cn/yzygj/c100068/ 202507/5b3f41180a42465eb9eec34597bacaf2.shtml, July 2025.